AI Pulse
📄 论文解读

让AI自己教自己看东西:用虚拟场景练出真实感知力

训练AI看图的常用办法是拿真人标注的数据喂它,又贵又慢。这篇论文换了个思路:让AI自己教自己。研究者用程序生成虚拟场景,场景里每个物体的位置、身份都是自动知道的,不需要人标注。然后让AI的“老师版”在回答问题时额外拿到一份“空间提示”——告诉它该看画面里的哪几个区域——老师因此能答得更准;接着让“学生版”只凭原图和问题去模仿老师的回答。学生没拿到提示,却学会了老师那种“知道该看哪儿”的能力。结果出人意料:虽然训练用的全是虚拟场景,但学到的能力迁移到了真实世界的测试上,在CVBench、V*、ZoomBench等6个真实感知基准上平均提升了3.23分,计数、图表、文档理解也都有进步。它不是你明天能用上的工具,但它指向一个更省钱的训练方向:用程序生成的场景替代昂贵的人工标注,让AI在自我对练中长出更广的感知力。

📄 原文摘要(英文)

On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains largely unexplored. Recent approaches use privileged visual information, such as image crops corresponding to a question, to improve fine-grained perception, but their gains are confined to tasks that benefit from such visual zooming and require either human-annotated grounding data or external teacher models. We introduce a different form of on-policy self-distillation for MLLMs that provides the teacher with textual, spatially grounded guidance identifying the visual elements relevant to a query. We use procedurally generated scenes with automatically available object identities and spatial coordinates, enabling scalable and annotation-free post-training. The teacher uses this spatial guidance to locate and integrate evidence from multiple relevant image regions, while the student learns to reproduce the resulting behavior from the image and question alone. Our approach consistently improves performance on counting, document and chart understanding benchmarks across multiple models. Importantly, although post-training uses only synthetic scenes, the resulting improvements transfer to real-world perception benchmarks, yielding a 3.23-point gain in average performance across CVBench, V*, ZoomBench, BLINK, HR-Bench, and MME-RealWorld. These results show that spatially grounded privileged information can induce broader perceptual capabilities through on-policy self-distillation, enabling substantial synthetic-to-real transfer beyond the task and data distribution used for post-training. Project page: https://github.com/sirkosophia/Where-OPD

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新