AI Pulse
📄 论文解读

让机器人看视频学动作,不再需要手把手教

机器人学动作一直卡在一个瓶颈:要教它干活,得先有人类亲手示范并标注动作,数据又贵又少。这篇论文换了个思路——让机器人直接看普通人拍的视频(比如人做饭、搬东西),从画面变化里自己悟出动作规律,再拿少量带标注的示范微调一下就能上手。它把「看视频学动作」和「执行动作」合并成同一个模型,不再需要先学视觉再转成动作的中间步骤。实验里,它用更少的标注数据就超过了老方法,还真的在实体机器人上跑通了。这不是你明天就能用的产品,但它指向一个更现实的方向:机器人或许能靠海量的日常视频长大,而不是只靠昂贵的实验室数据。

📄 原文摘要(英文)

World action models integrate future visual dynamics with robot action prediction, but their scalability remains limited by the need for action-annotated robot trajectories. Observation-only videos contain rich evidence about interaction dynamics, but existing approaches typically use them either to pretrain visual representations that must later be adapted for control, or to infer latent actions that are subsequently grounded to robot commands. We present NAVA-WAM, which introduces native action-prior learning by directly pretraining the action policy from observation-only videos, avoiding indirect representation-to-control transfer or a separate latent-action model. Our training consists of two stages. First, we pretrain on observation-only videos, where future-video flow-matching supervision over visual transitions is propagated through transition-structured joint attention to optimize the Action-DiT and learn action-relevant priors. Second, we use action-labeled demonstrations to post-train the Action-DiT for robot control through joint video--action flow matching, while asymmetric attention decouples the visual branch from iterative action denoising and enables efficient action-only inference. Extensive experiments show that NAVA-WAM consistently outperforms prior approaches under both in-distribution and out-of-distribution settings, while demonstrating strong action-label efficiency and effective real-robot generalization. These results establish native action-prior learning as an effective approach to directly pretrain action policies from observation-only videos, providing a scalable path beyond action-labeled robot data.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新