让机器人学会像人一样用手:先看人怎么做,再自己上手
现在的机器人抓东西,大多靠预设的抓取点或关节角度,遇到没见过的物体就抓瞎。这篇论文换了个思路:让机器人先看大量人类演示视频,学会预测「手和物体在 3D 空间里会怎么一起动」,再把预测出的手部动作翻译成自己的控制指令。关键是把世界拆成「场景」和「手」两部分,分别预测它们的 3D 轨迹,而不是像以前那样只盯着 2D 画面。在 10 个精细操作任务上,它比之前最好的方法平均高出 11.7 分,真实机器人上也跑赢了主流视觉语言模型。这不是你明天能用的技术,但它指向一个方向:机器人学操作,可能不再需要人类手写规则,而是靠「看懂人怎么动」来获得灵巧。
📄 原文摘要(英文)
World action models jointly learn to forecast world dynamics and predict robot actions, such that the learned internal world dynamics guide accurate actions. Existing approaches typically represent the world as RGB frames or latent counterparts while predicting actions as end-effector poses or joint angles, but they often struggle to capture the 3D spatial structure and contact geometry central to dexterous manipulation. We introduce Point World Action Model (PointWAM), a 3D world action model that decomposes the world into a scene (i.e., environment) and hands (i.e., actor), and jointly forecasts both as 3D point trajectories within a shared space-time coordinate frame. This explicit, disentangled representation enables effective pre-training on large-scale human demonstration videos without requiring any task-specific object or keypoint selection. Given a colored point cloud and a language instruction, PointWAM predicts how the scene and hands co-evolve in 3D space over time, then retargets the forecast hand motion to robot actions. Pre-training on human videos improves average DexJoCo success by 56.9 percentage points, and scene-trajectory supervision adds 10.9 points over forecasting the hands alone. With both, PointWAM surpasses the prior state of the art on ten DexJoCo tasks by 11.7 points and outperforms strong VLAs on a real robot.