AI Pulse
📄 论文解读

让AI画图,反而更懂看图

我们总以为AI画图和看图是两回事;这篇发现,让AI先学会画,它看东西反而更准。研究者把19种画图任务(比如预测深度、拼图、指认物体)和25种看图能力(比如数数、认类别、判断远近)一一配对,发现画图训练确实能提升看图表现,而且画图数据越多提升越大。有些对应很直觉:练深度预测就提升3D空间推理,练拼图就提升2D顺序感。但还有意外收获:练2.5D分割居然提升物体识别,练Z轴深度预测居然提升定位能力。他们用梯度对齐解释了这种迁移:画图和看图任务在模型内部越对齐,提升越大。这不是你明天能用的技巧,但它给了一个方向:想让AI更懂世界,也许该先让它动手画,而不是只喂它文字。

📄 原文摘要(英文)

Training a model to generate visual content can encourage it to learn rich perceptual capabilities related to geometry, spatial relationships, and objectness; yet, its benefits for visual understanding remain unclear. We ask: when and how does visual generation supervision improve visual understanding? We study controlled pairs of image-to-image (I2I) generation and image-to-text (I2T) understanding tasks that express the same underlying problem in different output modalities. We find that under the correct recipe, I2I training improves downstream I2T performance, with larger gains as the amount of I2I training data increases. We next ask which generation tasks benefit which understanding capabilities. To study transfer beyond paired tasks, we introduce OmniTaskonomy, a unified taxonomy spanning 19 I2I generation tasks and 25 I2T understanding capabilities. The resulting transfer map reveals selective, task-dependent benefits. Some follow intuitive correspondences, e.g., depth prediction improving metric 3D reasoning, object pointing improving counting, and jigsaw reconstruction improving 2D ordering. Interestingly, we also uncover surprising connections: 2.5D segmentation improving category recognition and Z-depth prediction improving localization. To probe these patterns, we analyze gradient alignment between generation and understanding tasks and find that stronger alignment is associated with larger downstream transfer gains. Together, our results highlight visual generation as a rich source of supervision for visual understanding and provide a roadmap for unlocking its benefits through the right training curriculum and task selection. Project page: https://omni-taskonomy.github.io/.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新