让AI用画图来思考,而不是只会写字
大模型现在会“想”了,但它的“想”全在文字里打转。这篇论文让AI把思考过程画出来:不是调用固定的视觉工具,而是直接让图像生成模型当“草稿纸”,用自然语言指挥它画图、改图,再基于画出来的东西继续推理。比如给它几张房间不同角度的照片,它能生成一张俯视图来理解空间关系;预测碰撞时,它先画出物体运动轨迹再判断。在六类视觉推理任务上,这种“画着思考”的方式比纯文字推理和专用视觉工具都强,最高提升25%。它不是你明天能用上的东西,但指向一个方向:AI的推理不该被锁死在文字里,画图本身就是一种思考方式。
📄 原文摘要(英文)
Chain-of-thought reasoning has revolutionized natural language processing by enabling large language models (LLMs) to decompose problems into intermediate steps before answering. Yet confining reasoning to the textual domain presents limitations for tasks requiring direct manipulation of visual representations. Recent efforts augment multimodal LLMs with external visual expert tools such as depth estimation or object detection modules, but these remain fundamentally limited by their reliance on narrow, rigid operations that cannot flexibly generate or transform visual content. We propose ReImaGin, which leverages image generation models as a flexible visual reasoning mechanism for multimodal LLMs: unlike fixed-function tools, they accept natural language commands and can perform open-ended visual operations, like removing an occlusion or generating a floorplan from multiple disjoint views of a room. Across six diverse visual reasoning tasks including multi-view spatial reasoning and collision prediction, ReImaGin consistently outperforms both text-only reasoning and specialist vision-tool baselines, with gains of up to 25\%, demonstrating the advantage of flexible, generative visual reasoning.