AI Pulse
📄 论文解读

给机器人造数据:把家变成录音棚

机器人学不会做家务,不是因为算法笨,而是因为缺数据——它需要同时看到手怎么动、脚怎么走、东西怎么被拿起、声音和触感怎么变。这篇把普通家庭改造成同步录制的「数据工厂」:桌上一个尺度拍手部精细操作,房间一个尺度拍全身动作,再加上第一视角、多视角视频、物体轨迹、音频和触觉,全部对齐成一条流。他们用这套设备录了150小时、7.5万个交互片段,覆盖200种家务任务。结果发现,现在最强的模型在接触、遮挡、自身运动和长时间跨度上仍然大量翻车。它不是你明天能用上的东西,但它是让机器人真正学会「过日子」所缺的那块拼图。

📄 原文摘要(英文)

Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新