AI画图翻车,现在让AI自己修自己的流程
多参考图生成一直有个尴尬:你给AI几张参考图让它合成一张,它要么漏掉某个主体,要么把两个主体生硬地拼在一起。过去人们靠手写流程代码来调,但效果全凭运气,不同人写的流程差距巨大。这篇论文干脆让AI自己改自己的流程:一个编码智能体反复重写那段控制代码,把「用来提建议的任务」和「用来选结果的任务」分开,再从表现最好的几个流程里继续搜。结果,开源模型FLUX.2在四参考图任务上从5.72分涨到7.37分,追平甚至超过Nano Banana Pro和GPT-Image-1.5这些闭源模型,而且换生成器、换参考图数量、换评测标准,这套流程不用重新优化照样有效。它不是你明天能直接用的功能,但它指向一个趋势:AI调自己的门槛正在降低,以后「调模型」可能不再是工程师的专利。
📄 原文摘要(英文)
Recent image generation models can take multiple reference images as input and combine them into a new image. However, multi-reference image generation remains challenging: models may omit or duplicate subjects from the references, or produce images in which multiple subjects appear unnaturally pasted. Recent work has proposed image generation agents that combine image generation models, reasoning models, and a harness, which is an executable program that specifies how reference images are interpreted, how generation is performed, how outputs are diagnosed, and how the final image is selected. In multi-reference generation, however, references play different roles and outputs must satisfy many criteria at once, such as fidelity to each reference and the naturalness of the whole image, so many parts of the harness could be improved, from how references are processed to how outputs are diagnosed. This makes it hard to predict which changes will improve performance and by how much, and good harnesses difficult to design by hand; indeed, human-written harnesses vary widely in performance. We therefore propose AutoRef, which optimizes the harness automatically while keeping both models frozen: a coding agent iteratively rewrites the harness code. AutoRef separates the tasks whose feedback informs proposals from the tasks used to select candidates, and continues the search from a beam of the top-ranked harnesses on the selection tasks. Using this procedure, we discover AutoRef-Harness, which improves the open-weight FLUX.2 [klein] 4B from 5.72 to 7.37 on held-out four-reference tasks of the MultiBanana benchmark, matching or exceeding proprietary models including Nano Banana Pro and GPT-Image-1.5. Without re-optimization, the same harness also improves results when the generator, number of references, benchmark, evaluator, or reasoning model differs from those used in the search.