AI Pulse
📄 论文解读

机器人看视频学动作,数据翻百倍成功率翻三倍

机器人学动作,过去靠人遥控演示,现在有个新路子:让它看大量视频自己悟。GE-Act 2.0 把「预测画面」和「反推动作」拆成两步,先看视频学世界怎么变,再学怎么动手,最后用「只挑和真实动作对得上的预测」来校准。结果:训练数据从 300 小时加到 3 万小时,零样本成功率从 17% 涨到 44%;而且它只用了 2% 的跨形态数据,就带来了 17.7 个点的提升,说明看别的机器人干活也能学。这不是你明天能用的东西,但它指向一个趋势:机器人可能不再需要人类手把手教,看够了视频就能自己上手。

📄 原文摘要(英文)

World-action models (WAM) predict future states to guide robot actions, enabling learning from both action-free video and action-labeled interaction. Most inherit pretrained video generators, leaving WAM pretraining and scaling underexplored. We introduce Genie Envisioner Act 2.0 (GE-Act 2.0), a world-action model whose trainable generative and action components are all initialized from scratch on manipulation data. It combines a control-oriented autoencoder (CoAE), a single-step visual planner (SVP), and an inverse dynamics model (IDM). CoAE retains action- and instruction-relevant information under aggressive compression, while SVP produces a complete future state in one differentiable pass, so visual planning and inverse dynamics can be pretrained separately on complementary data. The components are then jointly trained with knowledge-aligned selective optimization (KASO), which reduces mismatched supervision by selecting only predicted futures judged behaviorally compatible with the recorded action. We evaluate pretrained checkpoints directly, without per-task fine-tuning, on 100 tasks across 20 manipulation skill groups with held-out scenes, backgrounds, lighting, and object instances. Scaling co-training data from 300 to 30,000 hours raises success from 17.1% to 44.1% on G1-OP and from 13.4% to 31.1% on G2-90D; despite comprising less than 2% of the co-training data, G2-90D improves by 17.7 points, suggesting cross-embodiment transfer. Gains span 19/20 and 18/20 skill groups, and skill-specific coverage strongly correlates with zero-shot out-of-distribution (OOD) success (Pearson r=0.80; Spearman rho=0.85). Under the same protocol, the model grounds object, color, shape, and position references in at least 90% of trials and follows explicit instructions even when they conflict with an already-committed behavior or a conventional scene association.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新