图像编辑的瓶颈不在模型,在概念太粗
现在的 AI 修图,你说「把猫变成狗」它行,但你说「把这只猫的坐姿改成趴着、同时保留它的毛色和背景光线」,它就懵了。问题不在模型不够聪明,而在训练时喂给它的「编辑概念」太粗——只有「换猫」「换背景」这种大方向,没有「坐姿」「毛色」「光线」这种细颗粒度。这篇论文干了两件事:先建了一个超过 1000 种细粒度编辑概念的分类体系,再据此造了 1200 万对高质量编辑样本,让模型在训练时能同时学多个互不干扰的概念,而不是一次只学一个。结果是在细粒度编辑上明显超过之前的方法。它不是你明天就能用的工具,但它指出了一个方向:AI 修图的下一步,不是更大的模型,而是更细的指令。
📄 原文摘要(英文)
Existing image editing frameworks predominantly follow the training paradigm of text-to-image diffusion models. However, extending this paradigm to image editing highlights two inherent discrepancies, specifically, the insufficient attention to edit concept granularity and the training inefficiency caused by sparse supervision signals. To address these issues, we establish a comprehensive hierarchical taxonomy featuring over 1,000 fine-grained edit concepts and build ConceptEdit-12M, a massive dataset of 12 million high-quality editing pairs via an improved synthesis framework. This library-driven approach effectively rectifies the distribution collapse of generated data while ensuring high data fidelity. Furthermore, we propose a dense supervision training strategy that synthesizes multiple non-interfering concepts into single image pairs. By providing richer learning signals, this strategy significantly enhances both training efficiency and overall model performance. Training results validate our strategy, significantly outperforming prior works. Finally, we present ConceptEdit-Bench, a granular evaluation suite designed to diagnose model capabilities across a vast array of real-world scenarios.