AI Pulse
📄 论文解读

让AI预测视频下一秒,不再靠猜而是靠推理

现在的多模态大模型看视频,擅长总结已经发生的,却不擅长预测接下来会发生什么——它们往往靠文字惯性瞎猜,而不是真正看懂画面里的因果链。这篇研究给AI配了一套“侦探工具”:先让它把当前画面里的关键状态记下来,再放大局部细节、回查关键帧,像拼图一样补出中间缺失的环节,最后用强化学习奖励它“推理对”而不是“猜中文字”。在专门的预测测试上,它明显超过了更大的模型。它不是你明天能用上的东西,但方向值得注意:AI对世界的理解,正在从“事后总结”转向“事前推演”。

📄 原文摘要(英文)

Multimodal Large Language Models (MLLMs) have demonstrated remarkable potential in video understanding, yet their reliance on retrospective summarization and text-centric priors often limits their ability to bridge unobserved causal transitions when applied to Video Event Prediction (VEP). To address this, we propose VepAgent, an agentic framework that integrates causal-transition reasoning with tool-augmented reinforcement learning (RL) for robust VEP. Unlike prior methods that passively project future trajectories from historical dependencies, our approach explicitly models the logical progression from terminal observed states to future events. Specifically, we first construct futurebench-4K, a high-quality chain-of-thought dataset for supervised fine-tuning (SFT) that effectively bridges the causal-logic gap by structuring the deduction of unobserved intermediate states. Subsequently, we develop a diagnostic tool library integrating state tracking, frame retrieval, and region magnification, enabling the agent to dynamically augment reasoning with external tools to recover missing spatio-temporal evidence and resolve visual ambiguities during inference. Moreover, we propose a composite reward mechanism that jointly optimizes prediction accuracy, causal coherence, and reliable prior, compelling the agent to rely on genuine visual grounding rather than superficial textual similarities. Extensive evaluations on FutureBench and NEPBench datasets demonstrate that our method achieves state-of-the-art performance, significantly outperforming larger MLLMs and validating the empirical effectiveness of our agentic, future-oriented reasoning paradigm.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新