AI 重建 3D 场景,终于学会让物体「卡」在一起
以前的 AI 从一张照片重建 3D 场景,每个物体都是独立画出来的,沙发和茶几各摆各的,经常悬空、穿模、互相挤在一起。这篇让 AI 先看邻居:生成每个物体时,明确参考周围物体的几何形状和物理关系,让它们像积木一样严丝合缝地卡进场景里。为此研究者造了一个 120 万场景的物理模拟数据集,专门标注物体间的接触关系。结果是,哪怕物体被遮挡,AI 也能猜出它该长什么样、该摆在哪,物理稳定性和生成质量都刷新了纪录。它不是你明天能用上的工具,但这是 3D 场景生成从「摆拍」走向「真搭」的关键一步。
📄 原文摘要(英文)
We propose Tetris3D, a generative framework for single-image 3D scene reconstruction that recovers objects which are physically and geometrically coherent as a scene. Existing methods often generate objects independently or couple them implicitly, providing limited guidance for ensuring fine-grained spatial compatibility between neighboring objects that interact with one another. To address this, we explicitly condition the generation of each object on the geometry of surrounding objects and their physical relationships, guiding its shape and pose to remain geometrically and physically plausible within the scene. Moreover, we introduce ComOb, a physics simulation-based dataset of 1.2M scenes featuring physical interactions across diverse object categories, with per-object meshes and pairwise physical relation annotations. Comprehensive experiments on synthetic and realworld scenes show that Tetris3D recovers coherent object shapes and poses even when interacting regions are occluded, and achieves state-of-the-art performance in both generation quality and physical stability.