AI Pulse
📄 论文解读

修图AI不再一步步想,直接算出参数

现在的AI修图是“边想边做”:先读图、再推理、再选工具、再填参数,一步接一步,慢且费内存。这篇把整个过程压成一步——直接根据你的指令和原图,生成一组修图参数,像用公式算出来一样。效果不输那些会“思考”的AI,但速度快了至少50倍,内存省了近一半。它不是你明天就能用的产品,但它指向一个趋势:AI做事不一定非要“像人一样思考”,直接算结果,可能更快更好。

📄 原文摘要(英文)

Tool-based image editing (image retouching) is commonly formulated with autoregressive multimodal large language models (MLLMs) that sequentially generate reasoning, tool selections, and parameter values. In this work, we present a novel approach to tool-based image editing by framing the task as a flow matching problem. We introduce FlowTool, a framework that directly models the distribution of high-quality tool parameters conditioned on the input image and user instruction using conditional rectified flow. FlowTool combines a vision-language model backbone for multimodal understanding with a Diffusion Transformer parameter generator that transforms Gaussian noise into an editing plan. We train FlowTool with a two-stage supervised flow-matching curriculum, followed by reward-based post-training. Across MMArt-Bench, FlowTool-Eval, ArtEdit-Bench, and MIT-Adobe5K, FlowTool achieves significantly stronger reference-based performance than specialized MLLM editing agents and proprietary MLLMs, while remaining competitive with proprietary models under reference-free evaluation. Moreover, FlowTool significantly improves inference efficiency, reducing latency by at least 50times while requiring nearly 2times less memory than the compared baselines. These results demonstrate that tool-based image editing can be effectively modeled as conditional generation over structured continuous editing parameters, without autoregressive reasoning.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新