长视频生成终于不漂了:把缓存也拉回训练轨道
现在的 AI 生成短视频很稳,但一拉长就露馅:颜色慢慢变、画面开始糊、动作越来越僵。之前大家靠“缓存关键帧”来救,但没人管这些缓存本身是不是还在模型熟悉的范围内——结果缓存自己先跑偏了,越救越乱。这篇的做法很直接:在生成时强制每一段都只按训练时的规矩来,缓存不回头去看更早的内容,让滚动窗口始终待在模型“有把握”的区域。于是原本只能生成几秒的模型,被硬生生拉长到分钟级,画面不漂、动作不塌,用户评测也认。它不是你明天就能装进剪辑软件的功能,但这是长视频生成从“能看几秒”走向“能看几分钟”的关键一步。
📄 原文摘要(英文)
Modern autoregressive (AR) video diffusion models excel at short-horizon video generation, yet generating long videos remains challenging due to drifting, where colors and textures shift, and motion dynamics decay. Existing works primarily rely on KV conditioning, which selects or modifies cached key-value (KV) entries to mitigate drifting. However, we observe that KV conditioning alone is insufficient as it assumes cached KV entries remain in-distribution. This assumption fails beyond the training horizon: nothing constrains the construction of KV entries during rollout, giving rise to the KV-provenance problem where cached entries themselves become out-of-distribution (OOD). To address this, we propose In-Distribution Forcing (ID-Forcing), a test-time framework that aligns both KV caching and KV conditioning with training configurations. Its key mechanism, self-caching, prevents OOD KV entries at their source. Each chunk is cached without attending to prior KV entry, keeping the rolling window exactly in-distribution. Consequently, ID-Forcing seamlessly extends short-horizon models to minute-scale video generation. Extensive evaluations show that our method remains competitive on standard video generation benchmark while substantially outperforming prior work in mitigating drifting, as validated by both our drift metrics and a user study.