让AI在游戏里练级,4轮顶过百万次
AI 学新技能,通常要在虚拟世界里反复试错几百万次,慢得像蜗牛。这篇论文换了个思路:把游戏引擎和 AI 渲染器拼在一起,造出既真实又能实时互动的世界,让 AI 像人玩游戏一样边看边操作边学。结果惊人:只练 4 轮,就顶得上传统方法几百万次的训练效果。关键是它把「世界规则」写成代码,新关卡随时能造,AI 还能把经验写成「攻略」传给下一代。这不是你明天能用的工具,但它指向一个方向:AI 不再靠蛮力刷题,而是像人一样在真实感世界里快速成长。
📄 原文摘要(英文)
Interactive virtual worlds allow agents to learn through exploration and interaction. What agents can learn is bounded by the environments they practice in, which must be faithful, with consistent state, rules, and dynamics, and realistic, with observations that follow the real-world visual distributions. Achieving both across diverse worlds remains a bottleneck. We introduce AgentGarten, a framework that couples simulators and game engines with a shared neural renderer to build real-time interactive environments. Its simulation backends maintain persistent world state and execute program-defined interaction rules, while the renderer generates visual observations from structured conditions exported through a common interface. To build the neural renderer, we adapt a pretrained video model to geometry conditions, distill it with our proposed Adversarial Forcing, and optimize inference for real-time interaction. Adversarial Forcing makes history prefilling differentiable through exact replay, so that losses on later predictions update how the renderer encodes prior observations, and adds real-data adversarial supervision to improve its visual quality. In AgentGarten, agents perceive the world through visual observations, interact with it in real time, and improve by distilling each round of experience into playbooks that subsequent agents inherit and refine. Our empirical study demonstrates a substantial gain in learning efficiency, with agents learning from just 4 rounds compared with millions for a conventional reinforcement learning counterpart. As new worlds can be written as code and rendered through the same interface, environments can scale in both number and difficulty alongside their agents, a step toward agents that keep evolving through interactive experience.