AI 世界模型学会边玩边记,长视频不再崩
现在的 AI 世界模型能生成几秒视频,但一拉长就崩,而且你只能控制它动,不能控制场景里是谁、发生了什么。这篇把控制拆成三层:画面长相、角色身份、动态事件,分开喂给模型;同时把历史画面压成一小撮记忆,让模型边看边记,不用把整段长视频从头到尾重算。结果:它既能实时响应你的操作,又能保持几分钟的一致性,比现有方法都稳。这不是你明天能用的产品,但它指向一个方向——AI 生成的世界,正在从「几秒的片段」变成「能陪你玩下去的场景」。
📄 原文摘要(英文)
Interactive world models require responding in real time to versatile controls and maintaining long-horizon consistency. However, modeling heterogeneous controls remains difficult, while explosive contexts and unstable distillation impede achieving both long-horizon consistency and real-time responsiveness. In this paper, we present WorldPlay2, an interactive world model that couples a factorized hybrid control interface with a co-design of compressed memory and stable distillation. 1) Our factorized hybrid control interface integrates frame-aligned action control with structured semantic control that explicitly disentangles scene appearance, character identity, and dynamic semantic events, thereby facilitating effective control learning. 2) To achieve efficient long-horizon modeling, we compress historical contexts into compact memory tokens shared by the autoregressive student and the bidirectional teacher. This design enables clip-wise, memory-conditioned score evaluation instead of jointly processing an entire long rollout, substantially reducing distillation overhead. 3) We further propose Stable Forcing, which initializes the autoregressive student via a few-step strategy and leverages full-rollout replay to preserve the quality of long-horizon rollouts, ensuring robust and stable distillation. Extensive experiments demonstrate the strong generalizability of our model and its superior performance compared to existing methods.