让AI换环境不迷路:逼它用世界知识做决策
大模型当智能体时,一换环境就抓瞎。常见解法是让它预测未来画面,但既费训练又容易错上加错。这篇的思路反过来:模型预训练时早就把世界知识装进脑子里了,问题不是缺知识,而是没被逼着用。他们发现,常规训练里每个状态只给一个目标,模型很容易偷懒,靠表面习惯蒙混过关。于是他们搞了个叫EVOKE的方法:固定环境和历史,把同一个候选动作放到不同目标下排序,逼着模型必须真正理解世界才能答对。结果在多个任务上,换新环境的泛化能力和数据效率都明显提升。它不是你明天就能用的工具,但给「让AI更靠谱地干活」指了个新方向:与其教它更多,不如逼它用已有的。
📄 原文摘要(英文)
Large language models (LLMs) are increasingly deployed as agents for multi-step decision-making, yet transfer poorly to unseen environments. World-model methods address this by training agents to predict future observations, at the cost of additional training and errors that compound when predictions are used for planning. However, for LLM agents operating in digital environments, much of this world knowledge is already internalized during pretraining, which shifts the problem from acquiring it to eliciting it. We argue that typical post-training provides little pressure for such elicitation, since supervision under a single goal at each visited state inadvertently drives policies to rely on superficial contextual habits. We introduce EVOKE, a post-training method that supplies this pressure through goal diversity at fixed states. Motivated by theory showing that an agent competent across diverse goals must encode a world model recoverable from its action preferences, EVOKE holds the environment state and interaction history fixed and ranks the same candidate actions under alternative goals, forcing action preferences to change, so that a policy relying on contextual habits or single-goal correlations cannot order them correctly. This implicitly elicits the policy's pretrained world knowledge to inform decisions. We evaluate EVOKE across diverse tasks in three backbones, demonstrating improved task performance, unseen environment generalization, and data efficiency. We further conduct controlled analyses to better understand what drives these gains. These findings offer a new perspective on eliciting internalized world knowledge for transferable action through direct decision supervision.