AI Pulse
📄 论文解读

机器人学数据:五层金字塔,缺哪层都学不会

机器人不像大模型能靠互联网数据自学,它需要“身体”数据——物理状态、动作、触觉。这篇论文把机器人数据来源排成五层金字塔:真实机器人数据(最准但最贵)、人演示数据(UMI)、第一/第三人称视频、仿真数据、通用图文数据。越往下越容易规模化,但离真实机器人越远。研究者分析了当前机器人基础模型如何混合这些数据,发现触觉数据、失败数据、跨机器人动作对齐仍是空白。它不是你明天能用上的,但如果你关心机器人为什么还那么笨,这就是答案。

📄 原文摘要(英文)

Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this work, we organize the embodied data ecosystem as a "pyramid" spanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data. We organize the pyramid around the tension between scalability and robot alignment, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelity. We then analyze recent embodied foundation models through the lens of their data recipes, examining how different sources are selected, aligned, and mixed during pretraining. For embodied brain models, vision-language-action models, and world-action models alike, we relate data composition to capabilities in perception, reasoning, planning, action generation, and world prediction. We close by discussing six open challenges: building large-scale tactile datasets, collecting failure and recovery data, developing scalable data-collection pipelines, aligning actions across embodiments, leveraging egocentric data for dexterous manipulation, and designing principled data recipes for robot learning. We hope this work paves the foundation for the design of next-generation embodied systems.

arXiv 原文

📬 订阅 AI Pulse

每天三次更新,不错过重要信号

▲ 回到顶部