AI Pulse
📄 论文解读

换脸换声终于能同步了:一个模型搞定视频里的身份替换

以前的视频换脸和换声是分开做的,各管各的,结果就是脸换了但声音对不上口型,或者声音换了但脸还是原来的。这篇论文把两件事塞进同一个模型里,一次搞定:你给一段视频、一张参考照片、一段参考声音,它就能把视频里那个人的脸和声音都换成参考的,但动作、场景、说话内容都不变。关键是它用了一个「先抹掉再重建」的训练技巧:先把真实视频里的脸和声音特征抹掉,再让模型照着原样重建,这样它就能学会「换」而不只是「复制」。最终效果是音画同步、身份保留都做得不错,而且生成速度从30步降到3步,能实时流式输出。它不是你明天就能用的工具,但这是「数字人视频」从各管各的走向统一模型的一个明确信号。

📄 原文摘要(英文)

Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source motion, scene, linguistic content, and audio-video timing. Existing methods use separately optimized models for the two modalities, making audio-visual consistency difficult to enforce. We present UniSwap, the first framework for streaming joint audio-visual identity replacement in talking videos. Given a source video, a reference image, and a reference voice clip, UniSwap transfers the reference appearance and vocal timbre within a single audio-visual diffusion transformer while preserving the source content and dynamics. To address the scarcity of aligned cross-identity training pairs, we introduce a swap-and-reconstruct pipeline that removes visual and vocal identity from real clips and uses the original clips as reconstruction targets. Starting from a bidirectional backbone, we progressively adapt the model through In-context Pretraining for joint replacement, Conditional Streaming Adaptation for block-causal KV-cached generation, and Efficient Self-forcing DMD for mitigating exposure bias and reducing sampling from 30 to 3 denoising steps per block. Efficient Multi-LoRA Switching enables the three DMD roles to share a single frozen backbone. Feature-RoPE Decomposition keeps cached positions within the training range, supporting stable long-form inference. Experiments demonstrate strong audio-visual synchronization, competitive identity preservation, efficient streaming, and stable long-form generation.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新