AI世界模型能边走路边出声了
现在的AI生成视频,基本是“哑巴”:画面能动,但声音要么没有、要么对不上。这篇把“世界模型”往前推了一步:你像玩游戏一样操控镜头移动,它实时生成720p画面,同时配上和环境同步的环境音、音乐、甚至人物说话声,而且声音和画面是严格对齐的。它把“镜头想往哪走”翻译成一条统一的3D轨迹,第一人称和第三人称都能控制,长镜头下声音也不断。它不是你明天就能用的工具,但“能走进去、有声音的世界”这个方向,是生成媒体从“看片”走向“进入”的关键一步。
📄 原文摘要(英文)
We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes, camera--character dynamics are learned from data without view-specific controllers. Discrete commands and continuous poses are mapped to a shared metric-scale relative 6-DoF trajectory, with dataset-level calibration preserving motion magnitude across heterogeneous data. To jointly learn audio-visual generation and trajectory control, we construct a complementary data engine and adopt progressive training followed by autoregressive post-training for long-horizon generation. Extensive evaluations show that \model achieves strong trajectory following and high visual quality on public world-model benchmarks, supporting both first- and third-person interaction across varied subjects, and maintaining synchronized environmental sound and speech over long-horizon generation.