AI 预测未来时,想得比看得更重要
让 AI 边行动边预测未来,一直是机器人领域的热门思路。但有个争议:预测未来时,要不要真的把未来画面生成出来?生成画面很贵,所以有人干脆跳过这步,只保留预测的「想法」。这篇论文发现,跳过生成确实快,但 AI 的泛化能力——换环境、少数据、学新任务——会明显变差。更关键的是,差距几乎全来自第一步:AI 只需要「准备」未来,不需要「看见」未来。于是他们提出 Simple-WAM,把未来建模压成一次前向传播,既保住泛化,速度又和跳过生成的版本一样快。它不是你明天就能用的技术,但它回答了一个根本问题:AI 的想象力,到底该花在「想」还是「看」上。
📄 原文摘要(英文)
World action models (WAMs) predict the future alongside actions during training. Due to the heavy computation cost of video denoising, whether the future must still be generated during inference is disputed: Explicit WAMs denoise it into clean frames along with every action chunk, whereas Latent WAMs discard it entirely for acceleration. We find that latent WAMs, despite matching explicit ones on in-distribution tasks, fail to retain the generalization benefits that originally motivated WAMs. To demonstrate this, we evaluate generalization along three axes: environmental perturbation, data efficiency, and task generalization. Controlled comparisons with a matched backbone, training data, and budget reveal consistent degradation across all three axes when the action expert no longer conditions on future representations. Further analysis shows that the gap arises almost entirely from the first denoising step: the benefit comes from preparing the future, not generating it. We therefore propose Simple-WAM, which simplifies future modeling into a single forward pass of fully noised video tokens and adapts the training-time noise schedule to this inference behavior. Across simulation and real-world tasks, Simple-WAM achieves the best of both worlds, leading explicit WAMs in generalization performance with efficiency comparable to Latent WAMs. Project Page: https://zrporz.github.io/Simple-WAM-Web/