AI 用旧对话记录重建世界,模拟器不再需要原系统
训练 AI 智能体通常需要能跑的真实环境,但很多系统已经下线或无法复现。这篇论文换了个思路:不重建可执行环境,而是让一个“世界模型智能体”直接扮演环境,靠历史交互记录来模拟。做法是把旧记录整理成一本“世界书”,包含环境规则、证据和行为知识;运行时,世界模型智能体结合这本书和持续更新的状态,推断每个动作会带来什么观察和后果。在九个环境里,它的下一轮观察准确率和长程交互一致性都超过传统提示式方法;更关键的是,用它的模拟环境训练出的智能体,回到真实环境里操作时,动作有效性更高。它不是你明天能用上的东西,但指向一个方向:当原系统不可得时,历史痕迹本身就能成为训练场。
📄 原文摘要(英文)
Realistic environment replicas are increasingly valuable for training and evaluating LLM agents, yet the original systems may be inaccessible or impractical to reproduce. We explore agentic language world modeling: rather than rebuilding an executable environment, a world model agent serves as the environment for a task agent and supports faithful and stateful simulation. We instantiate this paradigm with Trace2Env, a learning-free framework for settings where the original system is unavailable but historical interaction traces remain accessible. Trace2Env reconstructs these traces into a reusable environment worldbook containing environment schemas, grounded evidence, and induced behavioral knowledge. At runtime, the world model agent actively consults the worldbook together with persistent episodic state to infer each action's observation and lasting state effects. Across nine environments, Trace2Env improves both next-observation fidelity and long-horizon interaction consistency over conventional prompt-based LWMs. In multi-turn interaction, task agent actions generated against Trace2Env remain valid more often when replayed in the real environment, indicating that its simulated dynamics better preserve the consequences of earlier actions across successive turns. These results establish agentic language world modeling as an alternative direction for building realistic environment replicas without reconstructing the original executable system.