AI Pulse
📄 论文解读

一张照片塞进10个人,AI终于不搞混谁是谁了

现在的AI生成多人合影,经常把张三的脸贴到李四身上,或者直接复制粘贴同一张脸。这篇论文的解法很直接:先给每个人发一个专属编号,再画一张座位图,最后按图索骥——每个编号只负责自己座位上的那张脸。关键一招是,它不靠AI自己猜谁是谁,而是直接用标注好的脸部区域去监督,把复制粘贴的穿帮率从16.9%降到5.5%。它不是你明天就能用的工具,但告诉你一个趋势:当AI要处理多个实体时,先规划布局再填充细节,比端到端硬学靠谱得多。

📄 原文摘要(英文)

Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location, while training-time identity losses must establish correspondence among several noisy predicted faces. We introduce WithEveryone, a unified framework for generating group images up to ten reference identities. WithEveryone injects each selected identity as an addressed token, predicts a structured identity--layout plan, and renders the plan as a visual condition. Its key objective, Layout-Grounded ID Loss, uses annotated face regions to supervise the intended identities directly, avoiding unstable embedding-based face matching; ID Representation Forcing additionally trains a prediction for each identity before image synthesis. On an identity-disjoint benchmark, WithEveryone achieves the highest target-context identity similarity, improving face similarity from 0.462 for GPT-Image-2 to 0.499, while reducing copy-paste artifacts from 0.169 to 0.055. It further covers 97.3\% of the requested identities with a duplicate rate of only 2.8\%. These results show that explicit identity--layout grounding enables identity-preserving generation to scale to larger groups without relying on direct reference-face copying.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新