让AI记住“上次看到过什么”
现在的视频世界模型有个尴尬:它要么把看过的画面全存下来(内存爆炸),要么压缩成一个小状态(细节全丢)。这篇论文给了一个两全方案:一半的注意力层保留原始画面缓存,另一半用“循环记忆”压缩,关键是——压缩时把摄像机的视角也写进去。于是当镜头转回老地方,模型能按视角精准调出“上次这里长什么样”。在公开基准上,它比同配置的全注意力模型省约30%内存,还能在固定内存下无限流式看长视频。这不是你明天能用的工具,但它指向一个方向:AI对世界的记忆,正在从“存照片”变成“懂空间”。
📄 原文摘要(英文)
When a camera revisits a previously observed region, a video world model should reproduce what was there before. This requires both remembering past observations and retrieving the right one for the current viewpoint. Key-value caches preserve visual detail but grow with video length; recurrent memory is compact but compresses history into a fixed-size state, so individual past observations are no longer directly accessible. We introduce LOCI, a hybrid spatial-memory architecture that keeps both representations. In half of the transformer blocks, main attention keeps a key-value cache of past observations; in the other half, it is restricted to the current chunk and complemented by a recurrent linear-attention memory whose reads and writes are conditioned on projective camera geometry, so viewpoint enters both memory addressing and stored content. Recurrent readouts flow into subsequent cache-backed blocks and supply their queries with accumulated scene context. On the public MIND memory benchmark and on held-out recorded trajectories, LOCI reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model; with full history, it lowers peak memory at equal length by about 30% relative to full softmax. With a bounded bank of retained observations, it streams long videos at constant memory and remains more faithful than full softmax under the same budget.