AI Pulse
📄 论文解读

视频生成提速9倍:一步把低清变4K

高清视频生成贵在算力,业界常用「先低清、再精修」的省钱路子,但精修本身要跑好几步,又成了新瓶颈。这篇把精修压成一步:先用高清数据持续训练,再用强化学习调优,最后蒸馏成单步模型,直接把低清输出抬到4K。在2K分辨率下,它一步的效果超过所有外部精修器;在3840×2176下,比跑三步的LTX-2.3还强,延迟快了8.91倍。它不是你明天能直接用的工具,但「一步精修」意味着视频生成的成本结构可能被改写——高清不再必然等于慢。

📄 原文摘要(英文)

High-resolution video generation is expensive, as its cost grows rapidly with the number of spatiotemporal tokens. A practical alternative first generates a lower-resolution video and then applies a refiner, but conventional multi-step refinement introduces a second sampling bottleneck. We present SoL-Refiner, a one-step video refiner that transforms low-resolution model outputs into 4K videos with a single denoising step. Our three-stage recipe combines high-resolution continual training, reinforcement learning (RL) post-training, and a final one-step distillation. We introduce Refiner-Bench, a video refinement benchmark constructed from the outputs of different video generators, and use a shared-input protocol to compare refiners at approximately 2K output resolution. At 2K, the one-step SoL-Refiner outperforms all external refiners on the VBench and UniPercept averages, while at 3840!times!2176 it improves both metrics over the three-step LTX-2.3 Refiner. With the complete acceleration stack, SoL-Refiner achieves an 8.91times speedup in refinement latency over the same baseline in our 2K latency setting.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新