机器人看更久的历史,动作反而更准
机器人控制一直有个两难:想看清动作和进度,就得看更长的历史视频,但处理历史会拖慢反应。这篇的核心发现是:历史长不等于用得上。只有当视频基础模型是自回归预训练(即按时间顺序预测下一帧)时,长历史才真正转化为更高的成功率——在 RoboCasa 测试中,把历史从 0 秒拉到 19.2 秒,成功率从 63.3% 升到 78.7%;而用双向预训练(同时看前后文)的模型,加长历史毫无收益。作者还做了工程优化,让模型在消费级显卡上也能实时跑,每个动作块(含未来视频预测)只需 107.4 毫秒。最亮眼的是动态叠杯测试:它 20 次全成功,而 Pi0.5 和 Fast-WAM 一次都没成。这不是你明天能装进家里的东西,但它说明一个趋势:机器人要真正干活,得先学会像人一样「记住刚才发生了什么」,而不是只看当前一帧。
📄 原文摘要(英文)
Real-time robot control demands enough visual history to infer motion and task progress, but processing that history can delay action. We present Long-WAM, a model-system framework for scaling the context of causal world-action models under real-time control constraints. Our central finding is that access to history is not the same as using it: longer histories pay off far more when the video foundation is pretrained autoregressively (AR). We first learn causal prediction from robot and egocentric videos without action labels, then preserve this history-to-future structure during world-action adaptation. On RoboCasa GR-1, increasing context from 0.0 to 19.2 seconds raises success from 63.3% to 78.7%, whereas a bidirectionally pretrained initialization shows no net gain; robot-domain AR pretraining further raises peak success on GR-1 and LIBERO-Long. Long-WAM also achieves the best results among compared methods on LIBERO-Long, RoboTwin 2.0, and DOMINO. Streaming observation encoding, asynchronous execution, and hardware-specific acceleration enable deployment on RTX 5090, DGX Spark, and Jetson AGX Thor without dropping future prediction; on RTX 5090, each action chunk, including future-video latent prediction, takes 107.4 ms. Real-time deployment on Unitree G1 and YAM supports dynamic and long-horizon manipulation, including 95% success on dynamic cup stacking, where Pi0.5 and Fast-WAM succeed in none of 20 trials. As a memory-informed executor, Long-WAM also complements higher-level planning in composite tasks.