AI Pulse
📄 论文解读

AI 记不清自己查过什么,才是它答错的主因

多模态 AI 答错题,大家总怪它看/听/读得不够准;这篇把锅甩给了「记性」:它把视频、网页、音频的观察全堆在对话历史里,越查越乱,后面的决策被前面的噪音带偏。作者做了个对照实验:只换掉规划模块,性能掉得远比换掉感知模块狠——问题不在眼睛,在脑子。解法是把「对话流水账」换成一张证据台账:哪些还没查到、哪些已确认、哪些互相矛盾,每条观察先过一道批评者过滤,只把有用的记上账。结果在 OmniGAIA 上准确率 81.4%,单题成本只有 Gemini-3.1-Pro 的四成左右。它不是你明天能装进产品的现成零件,但给了一个明确信号:下一代多模态智能体,拼的是怎么记账,不是怎么感知。

📄 原文摘要(英文)

Omni-modal agents must seek evidence across video, audio, web pages, and computation to answer questions. Their main bottleneck is planning: noisy multimodal observations accumulate in conversation history and disrupt later decisions, while multimodal models have limited capacity for multi-step planning. Controlled backend replacements support this diagnosis: replacing the planner causes a much larger performance loss than replacing the perception backend. We present Omni-Decision, an omni-modal agent built on evidence-ledger planning: it replaces the growing dialogue history with an explicit evidence ledger that records what evidence is still missing, what has been confirmed, and where records conflict. A critic reads each noisy observation and passes only the usable content to the ledger, discarding the rest, so the planner works from a compact context throughout the task. Each run records the state, action, and verdict at every step, and supervised fine-tuning and decision-level reinforcement learning on these trajectories further improve the planner. Omni-Decision achieves state-of-the-art accuracy of 81.4% on OmniGAIA at approximately 43% of Gemini-3.1-Pro's cost per question, and 65.0% on WorldSense long-video understanding, level with the strongest end-to-end model.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新