把文字压缩成图,AI 终于学会只看该看的地方
大模型读长文档时,一个省算力的老办法是把文字渲染成图片再喂进去——但分辨率低了看不清,高了又浪费。这篇论文让模型自己决定:先用低分辨率扫全局,再对关键区域放大细看,像人翻书一样。训练数据里每条推理都标注了答案在第几页、哪个位置,模型因此学会定位重点。结果是在同等压缩率下,准确率从 57.5 提到 87.4,速度还快了近 3 倍。它不是你明天能用上的东西,但指向一个趋势:AI 处理长文本的瓶颈,正在从『读得完』变成『知道该读哪』。
📄 原文摘要(英文)
Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI spends tokens on irrelevant content. We introduce FocusVTC, which breaks this trade-off through adaptive resolution while preserving general multimodal capabilities. It combines compressed low-DPI global views with selective region enhancement, integrating enhanced views into ongoing reasoning. We construct 29.4K high-quality Reasoning-Evidence Localization (REL) chain-of-thought examples (REL-CoT) that link reasoning traces to page indices and bounding boxes. Multi-resolution REL supervised fine-tuning (REL-SFT) teaches the model to localize relevant regions, and Group Relative Policy Optimization learns when to enhance resolution and how to use the resulting observations, without a separate continual-pretraining stage. At 72 DPI on RULER v1, FocusVTC scores 87.4 at 2.9times input compression, including tool observations, versus 57.5 for Glyph at 3.0times input compression. It surpasses its text-input backbone on LongBench (56.40 versus 55.86), improves the MRCR macro-average by 13.91 points, and achieves a 51.19 macro-average on VTCBench. The MRCR latency evaluation also shows a 2.79times online end-to-end speedup over Text. General multimodal capabilities are preserved, with MMMU increasing from 65.12 to 66.73 and MME from 2424.02 to 2457.62.