GPT-6 Astra把视觉难题分成了两半
通用大模型正在吃掉计算机视觉的“软”任务,但“硬”任务还守得住。研究者拿 GPT-6 Astra 和另外五个前沿系统,在 34 项能力、55 个基准上跟专用模型和人类对比,发现一个清晰的分界:凡是需要语义理解、推理、识别物体是什么的,通用模型已经逼近甚至追平专用模型;凡是需要精确几何测量、忠实重建、逐帧稳定的密集预测、细粒度专业知识的,差距依然很大。加推理和专用工具能补一部分,但效果因任务而异。换句话说,AI 越来越会“看懂”一张图,但还不太会“量准”一张图。这不是你明天能用的东西,但它划出了视觉 AI 的下一个战场:精度和保真,而不是理解。
📄 原文摘要(英文)
Frontier general-purpose systems are rapidly expanding beyond visual understanding into capabilities traditionally handled by dedicated computer-vision models. As these capabilities expand, a central question for the computer-vision community is how far this reach extends, and what remains hard. We evaluate GPT-6 Astra alongside five frontier general-purpose AI systems across 34 capabilities and 55 benchmarks spanning nine areas of computer vision. We compare their performance with dedicated models and humans where suitable references are available. Astra demonstrates broad visual capability, with substantial gains over other frontier systems in visual and spatial reasoning and several forms of structured prediction. Across the state-of-the-art systems, a consistent pattern emerges. Capabilities involving semantic interpretation, reasoning, and object-centric prediction increasingly approach or reach available reference levels. In contrast, larger gaps remain when tasks require metric geometric accuracy, faithful reconstruction, temporally consistent dense prediction, or specialized fine-grained visual knowledge. Additional reasoning and specialist tools close selected gaps, but their benefits vary across capabilities. These results map a changing landscape of computer vision in which increasingly sophisticated visual tasks are accessible through a general-purpose interface, while precise and fidelity-sensitive perception remains an important frontier.