AI Pulse
📄 论文解读

给机器人装上耳朵:听声辨位,不再只靠眼睛

机器人一直靠眼睛认路,但人类其实还会听声辨位——听到鸟叫就知道树在左边,听到车声就知道路口在右边。这篇研究给机器人补上了这个短板:他们建了一个包含真实世界声音的评测集,让机器人同时用眼睛和耳朵理解场景,再教它根据声音导航。结果,加了听觉的机器人在空间感知上明显更强,导航能力也接近只靠视觉的传统方案。它不是你明天就能用上的产品,但这是机器人从「看得见」走向「听得懂」的关键一步。

📄 原文摘要(英文)

Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challenging for embodied agents. In particular, it is still unclear how to effectively evaluate and model spatial audio understanding in embodied settings. To address this gap, we introduce OmniEchoBench, a unified benchmark for spatial audio-visual perception and audio-vision-language navigation. OmniEchoBench comprises six tasks over 197 real-world spatial audio-visual scenes, 2,972 question-answer pairs, and 900 navigation samples with first-order ambisonics (FOA) audio collected from 30 real-world environments. To enable scalable training supervision, we develop a controllable rendering pipeline for spatial audio. It preserves geometric consistency among sound sources, visual observations, and agent trajectories. Building on this, we propose OmniEcho, a spatially aware omni-modal model. It introduces an FOA spatial encoder alongside a pretrained semantic audio pathway. Extensive experiments show that OmniEcho achieves state-of-the-art performance on spatial audio-visual perception. For our sound-guided navigation, OmniEcho reaches a performance level close to that of traditional vision-language navigation. These results demonstrate that spatial audio can serve as a valuable signal for embodied scene reasoning and navigation, while also highlighting fine-grained spatial localization and distance estimation as important open challenges.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新