世界模型好看但不可靠
现在的 AI 生成视频、3D 场景越来越像真的,但一碰就露馅:你让它走两步、改个东西,它就开始犯迷糊。这篇论文给「世界模型」做了第一次大规模体检——14 个视频模型、9 个空间系统、8 个具身模型,在 1,138 个视频提示、300 个空间场景、254 个具身测试里,发现它们普遍「好看但不可靠」:视频模型在长时间生成和回头重看时前后不一致;空间模型放东西的准确率最高才 70%,改东西的成功率 73%;具身模型在多步动作后记不住状态,对动作和物理规则的改变反应迟钝。它不是你明天能用上的东西,但提醒你:别被 AI 生成的画面骗了,它离「懂世界」还差得远。
📄 原文摘要(英文)
Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification. We introduce HappyWorld-Bench, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them. Our design is built on a hierarchical capability framework of six world capabilities (W1-W6), from generative construction to unified world modeling, instantiated across three independent evaluation tracks: video world models, spatial world models, and embodied world models. HappyWorld-Bench comprises 1,138 video prompts, 300 spatial scenes, and 254 embodied test cases. Across all three tracks, we build and operate HappyWorld-Arena to organize human A/B comparisons and derive model-level Elo ratings, which complement newly designed automated metrics that capture behavioral correctness. We evaluate 14 video world models, 9 spatial systems, and 8 embodied candidates under this unified framework. Results reveal remaining reliability gaps across all three tracks: video models exhibit reduced consistency during extended rollouts and revisits, spatial models achieve at best 70.14% placement accuracy and 73.33% edit execution, and embodied models struggle to preserve state across multi-step actions and respond precisely to altered action conditions and physical rules. These findings highlight the need to evaluate world models not only by visual quality, but also by state consistency and the correctness of their responses to actions and interventions.