让AI记住去过的地方:世界模型不再“失忆”
AI在虚拟世界里逛了一圈再回来,常常“失忆”——它不记得自己刚才见过什么,画面会变得陌生。这篇研究让AI像人一样,对去过的地方有长期记忆。做法是:训练时不再要求AI把每一帧都看一遍,而是让它在“重访”时只去检索和当前视角最相关的旧画面,拼成一张共享记忆卡。这样,无论间隔多久,它都能保持场景一致,而且计算成本几乎不涨——从100秒拉长到400秒,耗时只多12%。这不是你明天能用的功能,但它意味着未来的AI游戏、自动驾驶模拟器,能构建出真正连贯、可信的虚拟世界。
📄 原文摘要(英文)
Streaming world models should render a place consistently across repeated visits. Directly supervising such revisits requires training samples that capture both visits, often spanning minutes. Yet dense attention over the full span incurs quadratic costs, making long-span supervision expensive. Memorizon breaks this coupling: long spans are needed for supervision, but not for attention, since the two visits can share a forward pass without including every intervening frame. A training sample covers a span of any length but is scored only on its last k chunks. Instead of tokenizing the history before them, each scored chunk retrieves its own top-K latents by camera co-visibility, and the union of these requests forms a shared bank. The bank is bounded by kK, so the sequence stays bounded however long the span; at the shortest span the recipe is exactly conventional training. Adding the bank raises the cost of a step once; beyond that, a longer span costs little, and going from 100 to 400 s adds 12% to the step time. Against a sliding-window baseline, retrieval raises revisit consistency on every split, and a span long enough to reach the first visit of each return adds a further 24% to 30%, at some cost in image quality; beyond that span, more length no longer helps. Filling the bank from another episode lowers revisit correlation by 83%, so the model uses what it retrieves. Project page: https://tingtingliao.github.io/memorizon