机器人学动作,不再需要看未来视频
机器人操作一直有个死结:要精准抓取,得预判物体接下来怎么动,但真让它看未来视频再行动,又慢又贵。这篇论文换了个思路——不生成未来画面,而是把「未来变化」压缩成一种内部信号,直接喂给动作模块。他们把视觉预测、场景理解、4D几何和动作生成塞进一个混合架构,用 2 万小时开放数据训练,在仿真和真机上成绩都超过此前方法。它不是你明天能用上的东西,但「不靠想象未来、靠感知变化」这个思路,可能是机器人摆脱慢半拍的关键一步。
📄 原文摘要(英文)
World Action Models (WAMs) jointly model visual dynamics and action generation for generalist robot manipulation. A central challenge is to integrate priors from large-scale pretrained models---including visual dynamics, scene semantics, geometry, and motion---into a unified framework for robot action generation. We introduce InternW0-Δ, a unified WAM pretrained on a heterogeneous corpus that outperforms prior methods across simulation benchmarks and real-robot platforms. InternW0-Δ combines pretrained visual dynamics, scene-level semantics, 4D geometric and motion priors, and action generation within a Mixture-of-Transformers (MoT) framework. A pretrained video expert and an action expert interact under semantic guidance from a frozen VLM, while a pretrained 4D foundation model injects geometric and motion priors through training-only distillation. We further introduce Causal Imprint, which learns future-relevant scene changes from training-only future supervision and provides predictive representations directly to the action expert without future-video rollout at inference. For large-scale joint training, we construct a heterogeneous corpus of robot demonstrations, UMI data, egocentric human demonstrations, and Ego2Robot data, curated and aligned under a common state-action representation. The resulting corpus contains over 20K hours of processed training data, to our knowledge the largest open-source corpus of its kind. We pretrain InternW0-Δ on this corpus and demonstrate strong performance across simulation benchmarks and real-robot platforms. We will open source the training code, model weights, infrastructure, data-processing pipeline, and processed data where licenses permit. Project page: https://internrobotics.github.io/InternW0-Delta/