AI Pulse
📄 论文解读

视频生成提速11倍,画质不掉:靠的是给AI留一张“底稿”

视频生成模型慢,一大原因是它要在海量像素上“想象”每一帧。现在有个办法:先把视频压成很小的“压缩包”再让模型处理,能快很多,但压缩太狠画质就崩,而且模型会“认不出”压缩后的数据,得从头再练一遍,代价极高。这篇论文的做法是:保留原模型认识的那份“底稿”,只把压缩时丢掉的信息单独存成一份“补丁”,两者合起来喂给模型;同时让压缩器学着“站在生成模型的角度”去压缩,而不是只顾着还原画面。结果在Wan2.1这个14B参数的视频模型上,压缩率提升8倍,生成速度提升11.1倍,画质评测和压缩前持平。它不是你明天就能用的工具,但它指出了一个方向:给大模型提速,不一定要换掉它的“大脑”,给它配个更聪明的“速记员”就行。

📄 原文摘要(英文)

Highly compressed video autoencoders offer an effective way to accelerate video diffusion models, as the Diffusion Transformer (DiT) operates on far fewer tokens. However, such autoencoders are challenging to train, since a higher compression ratio degrades reconstruction quality and recovering it requires more channels, which is known to slow the convergence of the DiT. The compressed latent also differs from the one the DiT was trained on, so the pretrained DiT must be either retrained from scratch or adapted at considerable cost. Compressing the autoencoder the DiT was trained with appears to preserve compatibility, yet optimizing it for reconstruction alone still shifts the latent away from the distribution the DiT has learned. To address this, we propose Generation-Aware Latent Compression for Efficient Video Generation (GRACE), a two-stage framework that compresses a pretrained video autoencoder while keeping it compatible with the pretrained DiT. Specifically, we keep a frozen base latent from the pretrained encoder and learn a residual latent for the information lost under stronger compression, while aligning the compressed latent with the pretrained latent in the feature space of the frozen DiT so that the autoencoder is optimized for generation. We then adapt the DiT with lightweight fine-tuning and asymmetric denoising, where the base is denoised ahead of the residual. GRACE reduces the token count of Wan2.1-I2V-14B by 8x and its latency by 11.1x at 480x832x81, while matching the generation quality of the pretrained pipeline before compression on VBench.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新