AI Pulse
📄 论文解读

AI生成视频+音频,终于有人用人类口味来打分

现在的AI生成视频和声音,评测时用的是分开的指标:画面清晰度、音质、同步率各打各的分。但人看着觉得“不对劲”的地方,恰恰是这些指标抓不到的——比如画面和声音在语义上根本不搭,或者和文字提示对不上。这篇论文先建了一个10K规模的人类偏好数据集,让人直接对比两个生成结果哪个更好,再训练一个“裁判”模型,它先学有明显差距的样本建立判断框架,再用人类标注验证过的推理去处理那些难分高下的对比,最后把人类反馈拆成多个维度分别强化学习,给模型更细的奖励信号。结果是这个裁判比传统指标更准,用它来训练生成模型,输出质量也明显提升。它不是你明天就能用的工具,但这是AI生成从“指标好看”走向“人看着顺眼”的关键一步。

📄 原文摘要(英文)

Using reinforcement learning to post-train joint video-audio generation models requires a reward signal. Existing methods construct this reward by combining metrics for individual quality dimensions, including audio quality, visual fidelity, and synchronization. However, these metrics evaluate perceptual dimensions separately and fail to capture the overall semantic and temporal coherence among the text prompt, video, and audio that shapes human preferences. Optimizing models against these metrics encourages reward hacking, generating video-audio content that achieves high scores on these metrics yet appears incoherent or unfaithful to human viewers. To address this problem, we first construct a large-scale human-preference dataset VAPref-10K for joint video-audio generation, comprising 9K prompts and 10.3K fine-grained paired comparisons from open-source generation models. We also introduce the VA-Judger-Bench benchmark with both in-domain and out-of-domain model comparisons to evaluate whether reward models truly align with human preferences. We further propose VA-Judger, a chain-of-thought omni-reward model for joint video-audio generation. In particular, VA-Judger first learns from pairs with clear quality gaps to establish structured output and coarse preference discrimination, then distills reliable preference explanations for harder near-quality comparisons via rejection sampling verified against human annotations, and finally performs dimension-wise reinforcement learning that decomposes human feedback into individual quality dimensions for denser reward signals than a single binary preference label. Experiments show that VA-Judger outperforms metric baselines in predicting human preferences on both in-domain and out-of-domain evaluations. Using its human-aligned rewards for post-training audio-video generation model also yields significant improvements in generation quality.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新