AI 看视频,其实没看懂物理
现在的多模态 AI 看视频,能认出「这是水」,却不知道水在怎么流。研究者造了 8,000 个物理模拟视频(流体、固体、光学等 80 类),给每个视频配上精确的物理参数,拿去考主流 AI 模型:让它根据文字描述找对应视频,几乎全错;让它判断两个视频是不是同一物理场景,结果接近瞎猜。但有意思的是,如果从 AI 的视频特征里直接「读」物理量(比如流速、密度),轻量模型反而能读出不少——说明 AI 不是没存物理信息,而是没把信息和画面「对齐」起来。更反直觉的是,用物理视频去专门训练 AI 的图文匹配能力,检索变准了,但直接读物理量的能力反而下降:对齐和量化,此消彼长。最后,研究者用这些 AI 特征去检索参考视频,喂给视频生成模型,生成的物理真实感确实变好了。它不是你明天能用上的东西,但它指出了一个方向:想让 AI 真正理解世界,光让它看更多视频不够,得让它把「看到什么」和「物理上发生了什么」绑在一起。
📄 原文摘要(英文)
Physical fidelity has received increasing attention in world models and video generation, yet how video representations encode physical information remains less understood. We introduce the World Embedding Benchmark, comprising 8,000 controlled simulation cases from 80 families spanning fluid mechanics, solid mechanics, dynamics, and optics & electromagnetism. Each case pairs a rendered video with simulation-derived physical annotations, supporting three complementary tasks: text-video retrieval, physical-property regression, and multiple-choice video-description pair classification. We use these tasks to distinguish cross-modal physical alignment from the recoverability of quantitative physical information. Evaluated pre-trained omnimodal embedding models show weak retrieval and near-chance within-family pair classification, while lightweight probes recover useful physical information from frozen video embeddings. Continual contrastive training with physics-specific video-text pairs improves retrieval and pair classification but degrades physical-property regression, revealing a trade-off between alignment and quantitative information recoverability. Finally, we use the embeddings to retrieve reference videos for retrieval-augmented generation with MiniMax-H3. Retrieved references improve the physical fidelity of generated videos, with stronger retrieval models yielding larger gains in our experiments. Together, these findings highlight the need to evaluate physical alignment and property recoverability jointly, and demonstrate the utility of physical representations for improving video generation.