AI Pulse
📄 论文解读

把AI的“想象空间”换成几何空间,视频生成终于不穿帮了

现在的AI生成视频,画面漂亮但一转动镜头就露馅:物体变形、透视错乱。研究者发现根子不在模型不够强,而在AI的“内部语言”——它天生只记颜色和纹理,不记空间位置。他们干脆把AI的想象空间整个换成一套“几何母语”:这个空间里每个点同时编码外观、深度、相机位置和3D坐标,生成器用这套语言说话,画面质量和3D一致性同时大涨,镜头轨迹误差直接减半。这不是给你明天用的工具,但它指了个方向:想让AI真正理解世界,得先给它一套能描述世界的语言。

📄 原文摘要(英文)

We present a compact geometry-native latent space as a shared foundation for perception and generation. Visual generators can produce photorealistic frames without preserving a consistent 3D scene. We argue that this is not only a modeling problem but also a representation problem: generators typically evolve appearance-centric latents, while perception models recover geometry in a semantically rich space that encodes cross-view structure. Rather than adding geometry as another output, we reparameterize a geometry foundation model's features into a compact latent space for generation. We realize this shift with the geometry-native autoencoder (GAE), whose latent is jointly decodable to appearance, depth, cameras, and point maps. With this state, a standard conditional flow supports diverse generation tasks. In controlled comparisons that hold the generator and training protocol fixed, replacing the latent with GAE improves both visual quality and independently measured 3D coherence: FVD falls by 12.7% and 23.1% on RealEstate10K and DL3DV, and camera-trajectory error is halved on RealEstate10K. Together, these results show that the latent space is central to geometry-consistent generation and can serve as a shared interface between perception and generation.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新