AI Pulse
📄 论文解读

AI试衣终于能一次穿多件,还不用抠图

以前的AI试衣只能换一件单品,还得先抠图。这篇把试衣当成「理解+生成」任务:你给一张模特图、几张衣服照片(甚至别人穿着的街拍),它直接合成模特穿上那些衣服的效果,不用你手动画遮罩。背后是三个阶段的训练——先让模型大量看衣服图,再学怎么穿,最后用两个裁判模型(一个专看衣服细节,一个看整体效果)反复调优。结果:单件试衣效果最好,多件混搭也领先现有系统,还能顺便改姿势。它不是你明天就能用的App,但方向很明确:以后买衣服,拍张自己、选几件衣服,AI直接给你看上身效果。

📄 原文摘要(英文)

We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-purpose image editor, Oxygen-TryOn is fashion-native, built for try-on through a dedicated data engine and try-on-specific training. Given one or more reference items (clean product shots or in-the-wild worn-on photos) and a single target subject image, it synthesizes a photorealistic image of the subject wearing the items across virtually any fashion category. Prior systems handle a single garment category in a studio setting, and recent multi-reference methods remain garment-centric; in contrast, Oxygen-TryOn supports diverse items and scenarios, including full- and half-body views, a variable number of references, and free multi-item composition, while faithfully preserving both subject identity and item appearance. Instead of mask-based inpainting, we reformulate try-on as a multi-reference, understanding-driven generation task. We build a data engine that collects, manufactures, annotates, and filters high-quality try-on data at scale, and design a three-stage recipe of continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL). The RL stage uses a hybrid reward combining an in-house try-on reward model with a proprietary, rubric-guided general-purpose model, jointly supervising fine-grained consistency and instruction-level quality. It also follows general editing instructions (e.g., pose changes) in the same pass. Across public benchmarks and our in-house Oxygen-TryOn Bench, it achieves state-of-the-art consistency and realism on single-item try-on and leads on multi-item try-on, matching or surpassing both leading proprietary systems (Nano Banana Pro, GPT-Image-2, Seedream5 Lite) and open-source models (FLUX.2).

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新