AI Pulse
📄 论文解读

AI视频生成:关键帧越多,效果越差?

现在的AI视频生成工具(如Sora、Runway)支持用多张参考图(关键帧)来指导生成,但研究者发现:关键帧越多,模型越容易“翻车”。他们构建了首个专门评测这一能力的基准KeyFrame-Compass,包含386个样本,覆盖不同场景、密度和输入格式。测试9个主流模型后发现:模型在忠实还原关键帧和保持视频自然流畅之间存在明显矛盾——关键帧越密集,视频越僵硬、越容易出错。更糟的是,大多数开源模型甚至无法正确理解“故事板”输入(多张图按时间顺序排列)。这不是你明天能直接用的工具,但它揭示了当前AI视频生成的一个核心瓶颈:模型还没学会在“听指令”和“自由发挥”之间找到平衡。

📄 原文摘要(英文)

Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it remains unclear whether they can faithfully reproduce the prescribed keyframes while maintaining overall video quality. We present KeyFrame-Compass, the first comprehensive benchmark for evaluating keyframe-conditioned video generation. The benchmark contains 386 carefully curated samples spanning three application domains, two video structures, two prompt granularities, two conditioning formats, and four keyframe densities, enabling controlled analysis under diverse generation settings. We further introduce an automated evaluation framework that jointly measures keyframe execution and overall video quality. Specifically, we decompose keyframe execution into six complementary metrics covering presence, fidelity, temporal ordering, localization, persistence, and uniqueness, while assessing overall video quality through evidence-grounded MLLM judgments augmented with specialized perception models. Experiments on nine representative video generation systems reveal several fundamental limitations. Current models exhibit a clear trade-off between faithful keyframe execution and natural video synthesis. Their performance further degrades as keyframe constraints become denser and most open-source models also fail to interpret storyboard-grid inputs as temporally ordered keyframe sequences.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新