AI Pulse
📄 论文解读

AI 智能体边干活边升级,不再靠死记硬背

现在的 AI 智能体解决长任务时,只有最后一步才知道自己做得对不对。过去它想变聪明,只能把经验写成文字存进记忆,下次靠检索碰运气。这篇论文换了个思路:让智能体在干活的同时,直接把自己的成功经验训练进模型权重里,相当于边工作边进修。关键技巧是让一个冻结的旧版自己当老师,用事后视角教新版怎么走,避免单次尝试的随机性把模型搞乱。在三个真实任务环境里,它越做越顺,成功率提升,还能把经验迁移到没见过的场景。它不是你明天能装上的功能,但指向一个方向:AI 不再需要专门的训练阶段,部署即学习。

📄 原文摘要(英文)

A large language model (LLM) agent solves long-horizon tasks through many reasoning-action turns, with one verification signal at termination. Deployed agents face streams of related tasks, making their trajectories a natural resource for improvement. In-context adaptation agents store reflections, memories, or skills as text, so reuse depends on retrieving the right experience and on a frozen policy executing it. We study Online Agentic Test-Time Training (OaTTT), which trains the LLM's weights on its own execution trajectories during deployment. The agent executes each task once, in one pass over the stream, and the executed trajectory with its verification result is the only learning signal for weight updates that persist across tasks. Directly imitating or reinforcing the generated tokens of this single attempt destabilizes the policy. We introduce ASCENT (Agentic Self-distillation for Cross-task EvolutioN at Test-time), which instead self-distills verified experience. A stable version of the LLM, its frozen initial copy, receives the verified trajectory as privileged information and predicts next-token distributions along it with this hindsight. Distilling them into persistent LoRA fast weights updates the agent for later tasks, without an external reference solution or stronger teacher. By further removing invalid-action turns, ASCENT distills enhanced privileged experience for more efficient execution. We characterize its population target and the limits of sparse outcome selection. Across ALFWorld, WebShop, and AppWorld at varied model scales, ASCENT improves task success and interaction efficiency as experience accumulates, outperforms online adaptation methods, and transfers to held-out scenes, showing that an agent can consolidate verified experience into its weights without a separate training phase or memory retrieval. Project page: https://artificer-ai-lab.github.io/ASCENT

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新