AI Pulse
📄 论文解读

让AI自己看自己画的图,再自己改,效果涨了12分

现在的多模态AI能画图也能看图,按理说它应该能自己检查自己画的图哪里不对、改一版、再看一眼、再改。但难点在于:改完好不好,得等图渲染出来才知道,所以「反思的文字」和「改图的动作」必须放在同一个模型里一起学。以前的做法要么只练改图、要么只练挑错,大部分提升空间被浪费了。这篇论文让AI在一条完整的「画→看→改→再看」闭环里用强化学习训练:同一张初始图生成多个修改版本,互相比较哪个改法更成功,再把奖励同时回传给「挑错的话」和「改图的动作」。结果在图像生成基准上比传统微调高了12个百分点,而且这个能力能迁移到没训练过的其他基准上。它不是你明天就能拿来用的工具,但它指向一个方向:AI的自我修正能力,可能不是靠堆数据,而是靠让它在闭环里自己练出来。

📄 原文摘要(英文)

Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped. We introduce UMM-Reflection, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment. Unlike single-round editing or pipelines with an external critic, credit flows across rounds and to both roles of the same model, and no verifier is needed at inference. On BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, and the gains transfer to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none of which is used in training.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新