给AI装长期记忆,它才记得住你刚才去过哪
现在的AI生成视频,记忆一长就乱:前面走过的房间、看过的物体,后面就忘了,画面跟着崩。这篇给世界模型装了一套「空间记忆系统」,核心是让AI自己决定记什么、忘什么——把相似场景归成一堆,每堆只留关键帧,再根据你下一步的动作去翻对应的记忆,最后过滤掉不可靠的。在多个模型上测试,生成稳定性和空间一致性都明显提升。它不是你明天能用的功能,但这是AI从「会生成画面」走向「真记得自己生成过什么」的关键一步。
📄 原文摘要(英文)
Long-video generation and world models have shown strong potential for interactive entertainment and embodied simulation by predicting future observations conditioned on user actions and historical memory. However, as memory sequences grow longer and their structures become increasingly complex, managing long-range spatial context becomes increasingly challenging, calling for a more intelligent and systematic memory-management strategy. Building on the advancing spatial reasoning capabilities of multimodal large language models (MLLMs) and the broader vision of unified models, we propose Spatial Memory Intelligence (SMI), the first framework to systematically employ an understanding model for spatial-memory management in long-video world models. SMI introduces four coordinated atomic operations: spatial clustering, within-cluster sparsification, action-aware retrieval, and reliability-aware filtering. Extensive experiments across multiple baselines, benchmarks, and world-model backbones demonstrate the effectiveness and generalizability of SMI, achieving comprehensive improvements in memory sparsity, generation stability, and spatial consistency.