AI Pulse
📄 论文解读

写提示词的人要失业了?AI 自己当导演

现在的文生视频,你给一句「一只猫在窗边看雨」,AI 就真的只拍这一句。但 WanPE 这个 3970 亿参数的模型,干的是导演的活:它先把你的一句话拆成一场戏——镜头怎么推、光线怎么变、声音什么时候起,再逐镜头生成画面。训练数据是 105 万个真实视频,靠「倒着写剧本」的方式学会从成品反推分镜。效果上,用它生成的视频,人类偏好比直接用你的原话高出 10 到 50 分,30 秒长视频的提升尤其夸张。它不是你明天就能用的工具,但它在改变一件事:以后你给 AI 的不是「提示词」,而是「一句话梗概」,剩下的分镜、运镜、节奏,AI 自己来。

📄 原文摘要(英文)

Video generation begins in text space by authoring a cinematic screenplay, then materializes into pixels. As contemporary video generators scale to 30 seconds and faithfully follow complex conditions, the textual prompt largely directs the production, planning how actions, camera trajectories, lighting, and sound unfold across multi-shot sequences. In this paper, we present WanPE, a 397B-parameter prompt enhancement model trained on 1.05M real-world videos to master director-level cinematic planning. WanPE formulates shot-level cinematic plans via video-grounded reverse construction and employs Semantic-Consistency GRPO (SC-GRPO) to faithfully preserve user requirements across shots and over time. To benchmark this capability, we curate WanPEval, a human-annotated testbed covering durations from 5 to 30 seconds across varying intent granularities, supported by approximately 11K blind pairwise assessments. When powering Wan3.0's video generator, WanPE-397B boosts human preference over raw user prompts by 10.66-18.84 points at 5-15 seconds and by a dramatic 50.86 points in the 30-second arena. Ablation studies show that reverse construction demonstrates clear superiority over forward rewriting, while SC-GRPO robustly preserves semantic fidelity across model scales. Ultimately, WanPE leads all evaluated commercial offerings at 5-15 seconds and remains competitive with Seedance 2.5 at 30 seconds.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新