让AI自己决定记什么,不再靠外部插件
现在的AI聊天时记性差,得靠外部系统帮它管理上下文——记什么、忘什么、怎么压缩,都是别人替它定的。这篇把上下文直接变成AI能自由读写的一个文件,让模型自己学会哪些信息值得留、哪些可以丢。结果:在多个任务上,准确率最高提升11.4%,算力消耗反而降了21.5%;在24小时的多智能体协作任务里,提升幅度大了65%。更关键的是,它把「记什么」从外部规则变成了模型的内在能力,还能用自然语言指令直接调教,甚至通过强化学习让模型越用越会记。这不是你明天就能用上的功能,但它指向一个趋势:AI的「记忆」正在从外挂变成本能。
📄 原文摘要(英文)
We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files. Building CLMs zero-shot with existing models outperforms SOTA context management strategies across a variety of tasks: 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm task. Moreover, by shifting context management from external harness control to intrinsic model behavior, CLMs naturally enable both in-context and parametric learning of context-management strategies. We show that CLMs can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute. We also introduce an online reinforcement learning method for CLMs, improving Qwen3.5-9B performance on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs. Finally, we co-design Suffix Cache Reuse for CLM serving, further reducing server-side compute by 35% relative to standard SGLang at matched performance.