AI 世界模型:4步生成视频,还能回放时再修
你玩《我的世界》时按一下键盘,AI 要立刻生成下一帧画面——但现在的视频生成模型通常要跑几十步才能出一张图,根本来不及。这篇把生成步数压到 1 步,而且能同时响应键盘和鼠标的连续操作。
核心做法是「渐进式蒸馏」:先让一个双向模型(能看前后帧)学会生成视频,再一步步把它压缩成只能看过去、只能跑 1-4 步的前向模型。每一步都确保模型对键盘按键和鼠标移动的响应不丢。最后还留了个后门:游戏时用 1 步模型快速出画面,回放时可以用 4 步模型把之前粗糙的草稿重新修一遍,效果接近 4 步模型,但比从噪声重新生成快了 3 倍。
在《我的世界》和第一人称射击游戏上,这个 1 步模型在画面质量、动作对齐、按键响应精度上都超过了现有方法。
[信赖·硬收] 这不是你明天能用的工具——它解决的是「AI 实时生成游戏画面」这个前沿难题,离产品化还有距离。但如果你关心 AI 怎么从「画图」进化到「能交互的世界模拟器」,这篇是重要的一步。
📄 原文摘要(英文)
Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step video synthesis, extending it to interactive world models remains challenging, as discrete keyboard states and continuous mouse motion must remain aligned with temporally compressed latent chunks during causal training and autoregressive rollout. We introduce ForgeWM, a progressive framework that transforms a bidirectional action-conditioned video generator into efficient few-step world models through domain adaptation, teacher-forced causal training, causal consistency distillation, and on-policy distribution matching with a bidirectional teacher. The resulting budget-specialized students operate at steady-state denoising budgets of 1, 2, and 4 steps. ForgeWM further supports a dual-path deployment protocol combining latency-critical interaction with optional replay-time refinement, where the one-step student re-noises and refines its saved draft. On paired Minecraft trajectories, ForgeWM leads the evaluated systems in Imaging Quality, reference-aligned motion-profile agreement, action-sign accuracy, and mouse-control accuracy, while achieving the lowest reference LPIPS; the same four-stage recipe transfers to gamepad-controlled FPS gameplay. Replay-time refinement matches four-step reference quality while remaining roughly three times closer to the experienced trajectory than regeneration from noise. These results demonstrate ForgeWM's effectiveness for controllable few-step video generation.