AI Pulse
📄 论文解读

AI 终于能生成带声音的视频了,连口型都对得上

以前 AI 生成视频,画面是有了,声音却是后配的,口型对不上、音效和画面各演各的。这篇把声音和画面当成两条并行的生产线,中间用双向注意力机制互相盯着对齐,于是生成 5 秒视频时,连 44kHz 的语音、口型、环境音都是同步的。它还能从一张图直接生成带声音的视频,并自动提升到全高清。这不是你明天就能用的工具,但它意味着「AI 拍视频」从默片时代进入了有声时代——下一步,短视频、广告、游戏过场动画的制作成本会再降一截。

📄 原文摘要(英文)

We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (I2AV) modes; a built-in super-resolution model raises the output resolution to Full-HD (1920times1080). Building on the video generation capabilities of Kandinsky 5.0, Kandinsky 6.0 Video employs a dual-stream CrossDiT architecture that connects a pretrained video stream and a newly trained audio stream through bidirectional cross-attention for temporal and semantic alignment. Our continuous pretraining strategy first trains the audio stream from scratch on large-scale audio corpora and then trains both streams jointly on paired audio-video data while preserving unimodal fidelity; pretraining is followed by supervised fine-tuning, reinforcement-learning-based post-training, and distillation. In side-by-side human evaluation, Kandinsky 6.0 Video Pro clearly outperforms its predecessor, Kandinsky 5.0 Video Pro, and remains competitive with leading audio-video generation models, particularly in speech quality. To accelerate open research and deployment in multimedia generation, we release the code, model checkpoints, and diffusers integration under the MIT license.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新