让AI视频生成先想好结局再动手
现在的AI视频生成是走一步看一步:预测下一帧,再预测下一帧,像蒙眼走路。这篇论文反着来——先让模型看一眼终点帧,再倒推中间该发生什么。做法是把目标帧提前塞进生成循环里,用不对称注意力让终点只做向导、不被中间过程干扰;同时让每一帧的隐藏状态提前对齐未来的画面,相当于每一步都在朝终点校准。结果是:只用25%的训练量就超过了完整训练的旧方法,在需要推理的视觉任务上更稳。它不是你明天能用的工具,但指向一个更本质的方向:AI生成不该只是像素的接龙,而该是目标的达成。
📄 原文摘要(英文)
Autoregressive (AR) video models excel at causal generation, but their reliance on next-chunk prediction confines them to a short-sighted, reactive paradigm. This limitation is particularly consequential for reasoning-oriented generation, where achieving a target outcome through valid intermediate states matters more than local visual plausibility. To address this challenge, we propose Learning Prospective Reasoning with Autoregressive Video Models (ProAR), a novel framework that transforms autoregressive video generation into a goal-oriented reasoning process. ProAR introduces two key components: (1) To anchor generation to the long-range outcome, we integrate goal-frame prediction into the autoregressive loop via an asymmetric attention mask, enabling the predicted goal frame to guide the generation of intermediate states without being disrupted by them. (2) To guide short-range transitions, we introduce future representation self-alignment to encourage current hidden states to anticipate upcoming temporal dynamics. By leveraging teacher-forcing in AR training, we extract clean future representations in a single forward pass and align current representations with them using a lightweight, training-only predictor. Together, these two mechanisms seamlessly combine explicit, sparse target supervision with implicit, dense step-wise guidance, promoting coherent, goal-directed reasoning progress with modest computational cost. Experiments show that ProAR's complementary components consistently improve performance across diverse visual reasoning benchmarks. The framework proves highly training-efficient, surpassing fully trained standard AR baselines using only 25% of the training steps. This paradigm also demonstrates promising applicability to embodied reasoning tasks.