让机器人看第一视角视频学干活
机器人学干活最缺的不是算法,是数据——真人示范太贵,但网上有海量第一视角视频(戴GoPro做饭、修东西那种)。问题是这些视频里镜头一直在晃,机器人分不清哪些动作是手在干活、哪些只是人脑袋在动。这篇把两者拆开,只提取「手和物体互动」的部分,再发现一个规律:人和机器人做同一件事时,慢节奏的动作模式是相通的,快节奏的各有各的样。于是它只拿慢的那部分去指导机器人,成功率在模拟环境里接近满分,真实世界四个任务也都能上手。它不是你明天就能用的产品,但方向很实在:以后机器人学新技能,可能不用再雇人录数据,直接看YouTube就够了。
📄 原文摘要(英文)
Learning general-purpose robot policies requires large-scale real-world interaction data, yet collecting robot demonstrations remains expensive and difficult to scale. Egocentric videos offer abundant human interaction experience with task-relevant semantics for robotic manipulation, but direct transfer is challenging for two reasons: latent actions inferred from frame reconstruction can be dominated by nuisance variation such as ego-camera motion, and human and robot behaviors often exhibit different temporal dynamics. We propose WING (World Action Learning via INteraction-Centric Spectral Latent Guidance), a framework for transferring interaction knowledge from egocentric videos to robot policies. WING first separates observer-induced motion from hand-object interaction and distills the interaction-centric component into latent actions. It then exploits the observation that cross-embodiment task semantics are concentrated in slowly varying temporal structures, identifying shared low-frequency components between egocentric latent actions and robot behaviors in the spectral domain and using them to guide action generation. WING achieves average success rates of 99.20% on LIBERO, 93.80% on RoboTwin 2.0, and 57.7% on RoboCasa-GR1, and also performs strongly across four real-world manipulation tasks under diverse generalization settings. These results show that interaction-centric spectral guidance provides an effective and scalable way to transfer physical interaction knowledge from human egocentric video to robot control. Project page: https://mikuz12.github.io/wing/