AI Pulse
📄 论文解读

AI拍长视频终于记得住剧情了

现在的AI生成视频,单看几秒很惊艳,一拉长就露馅:角色衣服变、道具消失、剧情前后对不上。根子在于AI只盯着当前这一帧,没有一本「账」记录故事里谁在哪、发生了什么。StoryEngine给AI配了这本账:把「剧情该怎样」和「画面实际长啥样」分开管,每次拍新镜头前先查账、按剧情推演状态,再生成画面;发现画面和账对不上,就局部重拍,不让小错滚成雪球。结果是长故事的角色、场景、因果都稳住了,在自建基准上全面超过现有方法。它不是你明天就能用的工具,但这是AI从「会拍片段」走向「会讲故事」的关键一步。

📄 原文摘要(英文)

Despite recent progress in agentic multi-shot video generation, producing coherent and consistent long-form stories remains challenging. Existing agentic pipelines typically rely on textual shot plans or previously generated pixels, yet lack an explicit mechanism for propagating the consequences of story events and maintaining the video world state across shots. As a result, missing visual details may be reconstructed inaccurately, while visual drift may propagate across subsequent shots, undermining both narrative coherence and visual consistency. To address these challenges, we propose StoryEngine, a state-grounded agentic framework for video storytelling. StoryEngine establishes a separation between authoritative semantic plans and unreliable visual observations. Specifically, StoryEngine maintains a structured representation of entity placement and story-relevant states, and propagates event-induced changes to define the intended start and end states of each shot. To visually realize these states, StoryEngine constructs canonical references for recurring entities and environments, and compiles state and visual constraints into executable render plans. Meanwhile, to realize these states correctly, a bounded evaluation-guided repair loop further corrects local state inconsistencies. Together, these mechanisms preserve causal story progression and prevent local visual errors from propagating across shots. To comprehensively evaluate long-form storytelling, we construct a benchmark across diverse scenarios and visual styles, with metrics assessing storytelling quality, narrative coherence, and visual consistency. Experimental results demonstrate that StoryEngine consistently outperforms state-of-the-art methods across all evaluation dimensions, validating its effectiveness for coherent and consistent video storytelling.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新