AI Pulse
📄 论文解读

AI 终于开始懂物理:生成的 3D 世界会遵守重力

现在的 AI 生成 3D 场景,基本是「画」出来的——它不知道重力、不知道深度,换个角度看就穿帮。这篇论文让 AI 同时建模物理(重力场、纬度)、几何(深度)和外观(图像),生成的世界不仅看起来对,而且物理上一致:东西会往下掉、视角移动时远近关系不乱。它还能在生成未来帧时把物理规律一路推演下去,而不是每帧重新瞎猜。这不是你明天能用的工具,但它指向一个关键拐点:AI 从「会画图」走向「懂世界」。

📄 原文摘要(英文)

We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules. To reliably construct and interact with 3D worlds, our framework jointly models three native world states: physics (gravity field and latitude), geometry (depth), and appearance (image), together with a unified Omni-Camera representation that supports diverse tasks and flexible motions. Beyond modeling these states, we introduce a strategy for propagating physical dynamics across future frames. By grounding absolute camera properties in the real world, Puffin-World enables physically consistent and visually stable world generation. We further couple appearance and geometry within a single generative process, jointly synthesizing each future view and reconstructing its underlying geometry. This unified paradigm enables interleaved closed-loop applications requiring synergy across multiple tasks, including mimic and self-calibrated world exploration. To scale Puffin-World to complex scenarios, we construct Puffin-16M, comprising 15 million vision-language-camera triplets and 1 million trajectories featuring various and challenging motions. To foster further research in this area, we released the code, models, and datasets.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新