视频生成模型会“反应”吗?新基准给出答案
现在的视频生成模型越来越被当作“世界模型”用,但现有评测只盯着画面好不好看、指令有没有照做。这篇论文提出一个四层基准,专门测一个被忽略的能力:模型能不能从画面状态推断世界该怎么反应,并生成输入里没明说的合理后果。测了20个主流模型,结果很分裂:靠镜头控制的模型镜头玩得溜,但没法动态交互;靠动作控制的模型能精确操控主体,但世界常常“没反应”;靠语言控制的模型交互好一点,但复杂指令又跟不上。没有一个模型能同时做到任务覆盖广和表现稳定——画面好看、指令听话,并不等于世界真的会“动起来”。
📄 原文摘要(英文)
Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inherent reactivity of the worlds they depict: the ability to infer from the scene state how the world should react and to generate plausible consequences not explicitly described in the input. Yet existing benchmarks mainly assess visual quality or explicit instruction fulfillment by checking whether requested actions and interaction outcomes are realized, leaving inherent reactivity underexamined. We introduce WorldExam, a hierarchical diagnostic benchmark spanning four levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity. It comprises 1,474 cases across eight dedicated tasks and supports unified evaluation of camera-, action-, and language-driven model paradigms. The World Reactivity level evaluates scene-conditioned reactions and goal-directed behaviors beyond what is explicitly specified in the input. Evaluation of 20 representative models reveals a clear capability split. Camera-driven models excel at camera control, but their interfaces do not support dynamic interaction; action-driven models control subjects more precisely but often leave the world unresponsive; and language-driven models perform better on interaction but follow complex controls less faithfully. No model combines broad task coverage with consistently strong performance, showing that high visual quality and explicit instruction fulfillment do not guarantee inherent reactivity.