AI Pulse
📄 论文解读

AI 拍片终于能听导演的话了

现在的 AI 生成视频+音频,画面和声音能对上,但镜头切换、台词开口的时间点,它根本不听你的。这篇提出 Temporal Context Routing:把剧本里写好的时间点直接映射到视频和音频的共享时间轴上,让每一句提示词精准落到该出现的位置。在 200 个测试剧本上,镜头切换的时间误差从 1.11 秒降到 0.042 秒,台词卡点准确率从 28.3% 提到 84.1%,画面质量和音画同步没掉。做短视频、广告、动画分镜的人,这是你明天就能拿去用的东西。

📄 原文摘要(英文)

Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchronization. However, they still provide limited control over when shot transitions occur and dialogue is spoken. This limitation constrains their application in script-driven content creation, where timing errors can undermine narrative coherence and the viewing experience. Current joint generators align video and audio representations on a shared temporal axis, yet the precise timing of shots and dialogue specified in a structured prompt is encoded only in the prompt's text representation and remains unaligned with the temporal coordinates of either modality. Consequently, video and audio may remain synchronized with each other while both fail to follow the script timeline. This mismatch motivates us to extend temporal alignment beyond video and audio to include the structured script. We therefore introduce Temporal Context Routing (TCR), which maps the script timing onto the shared temporal axis of video and audio generation and routes each prompt's guidance to the corresponding positions in both modalities. Compared with the baseline on 200 test scripts, TCR reduces Shot Boundary MAE by 96%, from 1.11 s to 0.042 s, and raises Dialogue [email protected] s from 28.3% to 84.1%. TCR achieves these improvements while maintaining visual quality and audio-visual synchronization comparable to those of the baselines. A user study further shows that participants prefer TCR on all five evaluated dimensions.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新