AI Pulse
📄 论文解读

AI 终于能对视频做深度研究了

过去 AI 的深度研究只擅长读文字和图片,一碰到视频就抓瞎:它没法把画面里的某个物体、人物关系、以及外部资料串成一条有依据的推理链。这篇把「看图找证据、查资料、组织答案」拆成一张统一的证据图,让同一个 AI 同时处理单图、多图和视频,在视频深度研究上比上一代模型强了 17.6 个百分点。它不是你明天就能用的工具,但意味着 AI 从「读文档」进化到「看世界」又近了一步。

📄 原文摘要(英文)

Single-image, multi-image, and video deep research require different visual operations but share a workflow of visual grounding, external retrieval, and fact composition. A key challenge is to preserve the dependencies linking localized visual anchors, entity relations, source-supported facts, and answer-producing operations. We introduce OneSearch-VL, a unified agent centered on the Visually Grounded Evidence Graph (VGEG), which encodes these dependencies as a shared task-level reference for data construction, process supervision, and operation-level evaluation. Our VGEG-based data engine constructs and verifies multi-image and video questions and filters expert trajectories. Using these data, we assemble OneSearch-VL-SFT-110K and OneSearch-VL-RL-10K for SFT and RL, respectively. We further derive the Evidence-aware Visual-Grounded Rubric reward (EVGR) from VGEG annotations to supervise evidence traceability and visual grounding during RL. For fine-grained evaluation, we construct OneSearch-MI-Bench and OneSearch-Video-Bench, organizing questions by the research operations encoded in their VGEGs. Experiments show that OneSearch-VL-8B improves over Qwen3-VL-8B with tool access by 20.2 and 17.6 percentage points on the two new benchmarks, respectively, while also achieving substantial gains across 7 image benchmarks and VideoDR. Project repository: https://github.com/appletea233/OneSearch-VL

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新