AI Pulse
📄 论文解读

机器人学会“预演”动作,失败率从48%降到6%

机器人做动作前,先在脑子里“预演”一遍视频,这并不新鲜。新鲜的是,这篇让预演变得诚实:以前的模型只见过成功案例,预演时总把失败也演成成功;这篇用“反事实训练”故意喂它失败的动作轨迹,让它学会预演真实的物理后果——比如推歪、没抓住。结果,在AgiBot挑战赛上,它把人类评估的交互缺陷率从48.12%压到6.25%,动作跟随也达到当前最好。它不是你明天就能装进家用机器人的东西,但它是让机器人“想清楚再做”的关键一步。

📄 原文摘要(英文)

We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction. Training such a model on existing robot datasets faces two obstacles: imprecise calibration can impair action following, while limited coverage of unsuccessful interactions can bias predictions toward successful outcomes. To improve action following across embodiments, we render action trajectories into image-space conditions and introduce offline geometric calibration to align these conditions with the target videos. To broaden interaction coverage, we introduce counterfactual post-training, modifying recorded action trajectories and generating future videos under a wider range of actions and contact configurations. To provide feedback on these predictions without paired ground-truth futures, we construct a human-annotated video dataset covering robot, object, and interaction defects and use it to train an embodied video reward model. Its scores guide reinforcement-learning post-training toward more physically plausible interaction outcomes. On AgiBot, DreamTrue attains state-of-the-art action following, while reducing the human-assessed interaction defect rate from from 48.12% to 6.25%. Notably, our model ranks first in the world model track of the AgiBot World Challenge 2026. The project page can be found at https://brave-eai.github.io/DreamTrue.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新