AI 终于能生成带声音的视频了,连口型都对得上
以前 AI 生成视频,画面是有了,声音却是后配的,口型对不上、音效和画面各演各的。这篇把声音和画面当成两条并行的生产线,中间用双向注意力机制互相盯着对齐,于是生成 5 秒视频时,连 44kHz 的语音、口型、环境音都是同步的。它还能从一张图直接生成带声音的视频,并自动提升到全高清。这不是你明天就能用的工具,但它意味着「AI 拍视频」从默片时代进入了有声时代——下一步,短视频、广告、游戏过场动画的制作成本会再降一截。
📄 原文摘要(英文)
We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (I2AV) modes; a built-in super-resolution model raises the output resolution to Full-HD (1920times1080). Building on the video generation capabilities of Kandinsky 5.0, Kandinsky 6.0 Video employs a dual-stream CrossDiT architecture that connects a pretrained video stream and a newly trained audio stream through bidirectional cross-attention for temporal and semantic alignment. Our continuous pretraining strategy first trains the audio stream from scratch on large-scale audio corpora and then trains both streams jointly on paired audio-video data while preserving unimodal fidelity; pretraining is followed by supervised fine-tuning, reinforcement-learning-based post-training, and distillation. In side-by-side human evaluation, Kandinsky 6.0 Video Pro clearly outperforms its predecessor, Kandinsky 5.0 Video Pro, and remains competitive with leading audio-video generation models, particularly in speech quality. To accelerate open research and deployment in multimedia generation, we release the code, model checkpoints, and diffusers integration under the MIT license.