AI Pulse
📄 论文解读

让AI画图不再纠结用哪层特征

现在的AI画图,常常要先把一张图压缩成“特征”,再靠这些特征重新画出来。问题在于:该用哪一层的特征?浅层保留细节好,深层生成效果佳,过去只能二选一。这篇研究干脆不选了——训练时随机抽几层特征来用,让模型学会“不管你给我哪层,我都能画”。结果同一个解码器,既能还原细节、又能生成新图,生成质量还比专门调过的版本好了27%。它不是你明天就能用的工具,但指向一个更省事的方向:与其为每个任务精挑细选,不如让模型自己适应。

📄 原文摘要(英文)

Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared latent space for the generator and pixel decoder. This choice involves a trade-off. Shallower layers tend to preserve fine pixel details better, while deeper layers tend to yield better generation metrics. A fixed heuristic layer fusion therefore couples two stages that benefit from different information. We introduce FuseReg, which replaces heuristic feature selection with training over random subsets of encoder layers. We theoretically analyze the underlying mechanism: subset sampling explicitly penalizes sensitivity to cross-layer disagreement. On ImageNet-256 with DINOv3-L, a single FuseReg decoder reconstructs from full, sparse, and single-layer fusions without retraining, achieving higher PSNR than decoders specialized to fixed fusions. This flexibility also benefits generation: decoder replacement alone reduces unguided gFID by 27% with an unchanged RAEv2 DiT-XL generator. The same regularization principle extends to diffusion training, with joint regularization of both stages reducing unguided gFID by 29% on DiT-Base. These results show that training downstream models for layer-fusion robustness narrows the reconstruction-generation gap without modifying the pretrained encoder.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新