AI Pulse
📄 论文解读

AI 终于能记住你身后发生了什么

现在的视频世界模型有个死穴:镜头一转,画面外的东西就忘了。AlayaVista 把世界拆成两层——一层是 360 度全景的“大背景”,一层是你眼前这个视角的“小窗口”。镜头动的时候,大背景一直在后台更新,小窗口只负责把你看到的那块渲染出来。这样既不用每帧重算整个场景,也不会一转镜头就穿帮。它用 1318 小时 4K 全景视频训练,镜头怎么转都接得上。这不是你明天能玩上的东西,但它指向一个更实在的未来:AI 生成的视频不再是一段段孤立的片段,而是能让你在里面“逛”的连续世界。

📄 原文摘要(英文)

Interactive video world models must maintain broad scene context under camera motion while producing high-fidelity observations with low latency. Existing approaches face a representation trade-off: perspective models operate on local views and must preserve off-screen content over long rollouts, whereas broader spatial coverage is typically obtained by synthesizing full-sphere videos or constructing explicit 3D representations. Motivated by the complementary roles of global context and selective local acuity in visual perception, we present AlayaVista, a camera-controllable streaming video world model that decouples panoramic world evolution from perspective observation synthesis. Given a single perspective image, AlayaVista constructs a 360-degree scene prior using a pretrained panorama expansion model and then evolves the scene as a camera-conditioned panoramic latent state. A latent viewport renderer maps this state to the requested perspective video latents, while a perspective refiner restores details, suppresses artifacts, and performs super-resolution. To support efficient streaming, we adapt the panoramic generator to chunk-autoregressive generation and distill both panoramic generation and perspective refinement into few-step processes. To provide the supervision required by this design, we construct MUGEN, a large-scale real-world panoramic video dataset containing 1,318 hours of videos at resolutions of at least 4K, together with rich semantic and geometric annotations.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新