AI Pulse
📄 论文解读

AI视频生成:用5秒训练,生成几分钟长视频

现在的AI视频生成,训练时用的是真实视频片段,但生成时用的是自己之前生成的画面,两者不匹配,导致长视频容易崩。这篇论文提出一种新训练方法:让AI在训练时就模拟生成时的“自产自销”过程,但通过巧妙的双阶段设计,让未来的画面能反过来指导早期画面如何存储记忆。结果只用5秒的训练窗口,就能生成几分钟的长视频,且主体、背景、时间稳定性都更好。它不是你明天就能用的工具,但指明了长视频生成的一个关键突破方向。

📄 原文摘要(英文)

Recent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained on histories produced by its own rollout rather than ground-truth video contexts. This reduces exposure bias, but the historical key-value cache is still used by future frames only as frozen rollout state. As a result, future losses cannot supervise how earlier generated latents should be written into more useful keys and values for later video-latent generation. We call this the historical context-gradient gap. We propose Self Gradient Forcing (SGF), a two-pass training strategy that restores this missing supervision signal without backpropagating through the full serial rollout. Pass 1 performs a no-gradient autoregressive rollout matching inference and, at a sampled denoising exit step, records both the self-generated context and the noisy latents fed to the model. Pass 2 performs parallel context-gradient reconstruction for the recorded exit step. The generated context is used as stop-gradient clean-latent input, while the model recomputes the context KV representations and future-to-context causal attention. Thus, SGF provides the missing memory-writing supervision within the native autoregressive training objective, using losses on future video latents to train the model to encode context into more effective causal memory. Across extensive long-horizon frame-wise and chunk-wise experiments under different initializations, SGF achieves stronger native long-video extrapolation than Self Forcing, especially in subject identity, background/layout consistency, and temporal stability. Remarkably, using only a 5-second training window, SGF can extrapolate to videos lasting several minutes. Code and models will be released to advance research on autoregressive video generation.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新