AI Pulse
📄 论文解读

生成视频时,AI 该记住哪一帧?答案藏在未来里

做长视频的 AI 有个老毛病:生成到后面,前面几百帧它记不住,画面就开始崩。现在的做法是回头看——哪一帧和当前最像就留哪帧。这篇论文反着来:先预测一下接下来最需要什么信息,再回头挑历史帧。它不生成完整未来,只生成一小撮「前瞻令牌」当路标,拿它们去历史里捞有用的帧。结果是在 5 个基准、11 个生成模型上,长程一致性、画面质量和动作对齐都稳定变好。它不是你明天能直接用的东西,但方向值得记住:AI 的记忆不该由过去决定,该由它接下来要做什么决定。

📄 原文摘要(英文)

Long-horizon video generation requires models to effectively leverage an increasingly long generation history. As the generated history grows, retaining all previous content becomes increasingly expensive and redundant, making effective historical selection essential. Existing approaches often determine historical relevance based on the current content. However, information relevant to the present is not necessarily useful for future generation, while seemingly less relevant history may become important later. Our key insight is that historical information should be selected according to its relevance to future information needs. Capturing these needs does not require generating the full future; instead, a compact representation of what becomes important next is sufficient to guide historical selection. Building on this insight, we propose FrameMorrow, a prospective frame selector that predicts a small set of prospective tokens representing future information needs and uses them to identify relevant information from history. FrameMorrow selects explicit historical frames rather than model-specific internal states, enabling plug-and-play integration across diverse generators, including closed-source models, with little additional inference cost. We evaluate FrameMorrow across five benchmarks and 11 generative models spanning long-video generation, interactive generation, and action-conditioned world models. Extensive experiments demonstrate consistent improvements in long-range consistency, visual quality, and action alignment across diverse generation settings.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新