AI Pulse
📄 论文解读

2000小时人做家务视频,机器人学得会吗?

机器人学做家务,最缺的不是算法,而是人干活的数据。Open-AoE 直接拉500多人用手机拍了2000小时日常操作视频——开冰箱、拧瓶盖、叠衣服,还附带了手部姿势、相机轨迹和动作标签。它不只是个数据集,还配了一套工具链:从原始视频自动切出动作片段、标出语义、重建手和相机轨迹,最后直接喂给机器人模型训练。你明天用不上,但它是让机器人从“看人做”学会“自己做”的关键基础设施。

📄 原文摘要(英文)

Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resources rarely combine low-cost continuous capture, manipulation-level structured annotations, and reusable tools for robot learning. We present Open-AoE, an open, community-oriented egocentric manipulation dataset and toolchain spanning the full pipeline from smartphone capture to model training. Its first release contains approximately 2,000 hours of manipulation video collected in natural environments by 500+ contributors using 400+ smartphones. The dataset provides text annotations, MANO-based hand poses, camera trajectories, and temporally localized atomic actions. Open-AoE further includes a data processing pipeline that transforms raw recordings into structured samples through temporal action segmentation, semantic annotation, hand reconstruction, and camera trajectory reconstruction. Meanwhile, we provide a separate downstream toolchain supports visualization, cross-embodiment retargeting, model-specific data conversion, and training recipes for VLA policies, WAMs, and World Models. By integrating scalable capture, structured processing, and downstream adaptation, Open-AoE reduces the barriers to both data contribution and reuse, providing practical open infrastructure for embodied model training, human-to-robot transfer, and world modeling.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新