AI Pulse
📄 论文解读

把AI的“经验”做成树,学一次,到处用

现在的AI智能体学做多步任务,是把每一步都当独立动作来练,像背课文一样从头背到尾。这篇论文发现,真正的高手不是这么干的:人会先记住“怎么订机票”这种套路,再在具体任务里套用。研究者干脆把AI反复用到的动作片段自动抽出来,拼成一棵“经验树”,让AI直接在这棵树上训练。结果在三个测试环境里,成功率比传统方法高4到6个百分点,而且不靠大模型额外生成提示,纯靠数据自己长出来。它不是你明天就能用的功能,但指向一个方向:AI的“经验”不该是一次性的,而该是可复用、可积累的资产。

📄 原文摘要(英文)

Multi-step agents are trained on flat action streams: SFT and RLVR weight every token uniformly and ignore the sub-procedures that recur across tasks, the hierarchy that lets humans plan top-down from reusable routines. This structure sits unused, and flat training uses each scarce trajectory less fully than its content allows. Recent agents do use that structure, but only as LLM-written skills in context, never in the weights, so their gains do not generalize beyond retrieval. We instead recover this hierarchy from the data itself and train on it, with no LLM calls. Following text tokenizers, which build a vocabulary by counting alone, we score action spans by reusability and merge canonicalized actions into a reusable eXperience tree (X-Tree). Each X-Tree node captures how a frequent and success-bearing skill is composed from sub-skills, guiding efficient generalization. We integrate X-Tree into three training settings: offline RL, with each node as a training instance; online RLVR, with an adaptive skill bonus; and on-policy self-distillation, with X-Tree as the self-teacher's privileged context. Across WebArena, ScienceWorld, and WebShop at three model scales, X-Tree improves over standard recipes at matched data and budget by up to 4.5% SR on WebArena, 5.8% SR on ScienceWorld and 4.1% success on WebShop. Matched analyses attribute the gains to the X-Tree structure and the three integrations.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新