AI Pulse
📄 论文解读

让AI像人一样在脑子里转图,推理准确率暴涨39个百分点

现在的AI看图说话,基本靠语言硬猜:你问它一个物体转90度后长什么样,它不会真的在脑子里转一下,而是从训练数据里找文字规律。这篇论文给AI装了一个轻量的“世界模型”分支,让它先自己生成旋转后的画面,再基于这个画面作答。在2D和3D心理旋转任务上,这个做法比传统微调高出最多39个百分点;而且一旦把生成的画面删掉或弄脏,成绩立刻暴跌——说明它确实是靠“在脑子里转图”而不是蒙对的。它不是你明天能用上的东西,但这是AI从“会说话”走向“会想象”的一个实在信号。

📄 原文摘要(英文)

Humans often solve spatial problems by mentally simulating visual transformations. In contrast, conventional vision-language models (VLMs) reason primarily through language. We investigate whether VLMs can solve spatial problems by reasoning with both text and generated visual states. To this end, we introduce WM-VLM, which equips a pretrained VLM with a lightweight world model branch for generating intermediate visual states. Our two-stage training first teaches the model to generate the next visual state and then to use that state for reasoning. We programmatically construct spatial reasoning tasks with verifiable intermediate visual states. These tasks allow us to evaluate how well the model generates visual states and how much it relies on them to answer the question. On 2D and 3D mental rotation tasks, WM-VLM consistently outperforms the supervised fine-tuned backbone, with gains of up to 39.25 percentage points. Ablations suggest that these gains depend on the generated visual states, as removing or corrupting them sharply reduces performance. Together, these results suggest that internal world models offer a promising path toward VLMs that reason in both language and visual space.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新