AI Pulse
📄 论文解读

AI 视频生成当世界模拟器,连倒杯水都做不全

把视频生成模型当「世界模拟器」来用,是当下很热的方向:给它一张初始画面和一个目标,让它自己推演接下来会发生什么。但新基准 Ego2Act 给这个幻想泼了盆冷水——它用 110 个真实日常任务、2640 段第一人称视频来考模型,结果发现模型经常跳过步骤、做一半就停,后面的关键状态根本没出现,目标自然完不成;尤其在复杂物体操作和持续世界建模上,物理细节一塌糊涂。它不是你明天能用上的东西,但它划出了一条硬底线:现在的视频模型离「能替我们推演世界」还很远,连倒杯水这种多步操作都演不圆。

📄 原文摘要(英文)

Video generation models are increasingly being explored as world simulators for embodied planning and learning. To do so effectively, these models must not only generate visually appealing frames, but also predict how environments dynamically evolve when executing goal-directed actions. While evaluating these capabilities is crucial, existing benchmarks focus mainly on single short actions or step-by-step instructions. This leaves multi-step physical reasoning underexplored, especially in egocentric video generation that requires planning to simulate proper execution to accomplish high-level goals by carrying out multiple real-world manipulations. We introduce Ego2Act, a goal-directed benchmark featuring 2,640 videos from 110 real-world tasks across day-to-day settings, varying object clutter and multi-step complexity. Given an initial scene image and a high-level goal, Ego2Act evaluates whether video generation models can produce realistic egocentric videos of a hand manipulating objects to carry out the task. To support scalable evaluation, we also introduce Ego2ActJudge, a reference-free evaluation pipeline that achieves better task completion and physics plausibility evaluation alignment with human consensus compared to relevant baselines. Our findings reveal that models' generated simulations often skip or partially execute steps, leaving later steps missing dependent states, which leads to unfulfilled goal. Furthermore, models consistently fail at fine-grained physical dynamics, particularly during complex object manipulation and persistent world modeling. We hope Ego2Act provides a rigorous testbed for advancing video models toward physically plausible, goal-directed simulation.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新