世界模型终于有声音了,而且带空间感
现在的 AI 世界模型只会生成画面,是哑巴。这篇让画面和声音在实时交互里一起演化:你转动视角,声音也跟着变——左边有车,声音就在左耳;你转头,声场跟着转。做法是把相机轨迹和用户操作作为条件,先训练一个双向教师模型,再蒸馏成能单 GPU 24 帧实时跑的流式学生,同时配了一套评测基准来检验声音是否真的跟着视角走。它还不是你能玩上的产品,但这是世界模型从「无声电影」走向「真实世界」的关键一步。
📄 原文摘要(英文)
World simulation is inherently multisensory, demanding synchronized visual and acoustic dynamics in real time. Yet prevailing interactive world models remain strictly silent, focusing exclusively on visual rendering and control while overlooking the acoustic dimension. We present HelixWorld, a real-time interactive audio-visual world model where visual scenes and camera-grounded spatial stereo sound co-evolve natively under user interaction. We curate a high-fidelity spatial audio-visual dataset with true stereo acoustics and metric camera poses, upon which we pre-train a bidirectional teacher conditioned on 6-DoF camera trajectories and user actions. To enable low-latency causal interaction, we distill the teacher into a few-step streaming student via an online trajectory distillation loss, sustaining drift-free joint audio-visual rollouts at 24 FPS on a single GPU. Furthermore, we formalize spatial-acoustic consistency and introduce HelixBench to evaluate whether synthesized sound fields faithfully track dynamic viewpoint motion. Extensive experiments demonstrate that HelixWorld matches state-of-the-art silent world models in visual fidelity and responsiveness, while significantly surpassing existing baselines in camera-aligned spatial-acoustic immersion.