AI Pulse
📄 论文解读

AI视频生成终于记住自己看过什么了

现在的AI视频生成像金鱼:能生成几秒惊艳画面,但镜头一转,场景就悄悄变了样——墙纸换了、杯子消失了。WorldCrafter给视频模型装了一种「3D记忆」:它把之前看过的画面按空间位置存起来,你转动镜头时,它先查一下那个方向原来长什么样,再生成新画面。结果就是,你能在一个AI生成的场景里连续逛上一分钟,来回转头,东西还在原位。这不是你明天能用的工具,但它指向一个关键拐点:AI生成视频从「单次表演」走向「可探索的世界」。

📄 原文摘要(英文)

Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory for this purpose. The key insight is to let the requested viewpoint shape how multi-view evidence is compressed into the video generator's limited token budget. Trained jointly with the video generator, a memory encoder and pose-conditioned readout module integrate historical observations into a fixed set of target view-specific tokens before denoising, without explicit depth-based correspondences. By combining this memory with recent temporal context and few-step distillation, WorldCrafter enables streaming scene exploration from a single input image or text prompt. Experiments across static and dynamic scenes show substantial gains in long-horizon consistency and camera-control accuracy while preserving visual quality during minute-scale exploration.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新