AI Pulse
📄 论文解读

让AI视频里的物体真的会动,而不是贴图

现在的AI视频生成,你让镜头绕一圈,画面里的车可能就变形了——因为模型只是按像素猜运动,没有真正的3D几何。这篇论文换了个思路:先把画面里的每个物体重建成一个3D模型,然后你告诉它这个模型每帧怎么平移、怎么旋转,它再据此渲染出视频。物体不会凭空长出新结构,镜头怎么转都保持同一辆车。它比你之前见过的控制方式都精确,但代价是:你得先接受一个前提——画面里的东西得能被重建成立体模型。这不是你明天就能用的工具,但它指向一个明确的方向:AI视频正在从“画得好看”走向“摆得对”。

📄 原文摘要(英文)

Precise control over camera and object motion is essential for professional video production. Existing methods control objects only coarsely, through image-plane cues that are ambiguous in depth and rotation or through 3D tracks and blobs that lack complete geometry and lose consistency across viewpoint changes. We introduce 4Director, a video world model conditioned on an explicit 4D scene representation: each object is reconstructed once from the input image as a canonical mesh and moved by one prescribed rigid transformation per frame. This representation provides an intuitive 3D control interface and prevents unobserved geometry from being regenerated independently in every frame. We render the controlled scene as a depth video and introduce a Motion Adapter that transforms this geometric scaffold into video while synthesizing view-consistent appearance, illumination, and non-rigid dynamics. For training, we construct RealCOD-Rigid, a new dataset of 20,774 clips annotated with rigid 3D scenes by our automatic pipeline. We further introduce Identity-Gated IoU (IG-IoU), which jointly evaluates adherence to prescribed object motion and preservation of object identity. Experiments demonstrate that 4Director consistently outperforms prior methods in visual quality and in camera and object control.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新