Research & Papers

Oxygen-TryOn achieves SOTA in any-item virtual try-on, beating GPT-Image-2 and FLUX.2

Fashion-native foundation model handles multi-item, full-body try-on with unprecedented realism.

Deep Dive

Oxygen-TryOn redefines virtual try-on by treating it as a multi-reference, understanding-driven generation task rather than mask-based inpainting. The model accepts one or more reference items (clean product shots or worn photos) and a single target subject image to produce a photorealistic output. It covers diverse fashion categories, supports variable numbers of references, and can follow general editing instructions such as pose changes. The system is built from scratch for try-on, powered by a large-scale data engine that collects, manufactures, annotates, and filters high-quality data.

Training follows a three-stage recipe: continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL). The RL stage uses a hybrid reward combining an in-house try-on reward model with a proprietary rubric-guided general-purpose model, jointly supervising fine-grained consistency and instruction-level quality. On the Oxygen-TryOn Bench and public benchmarks, Oxygen-TryOn achieves state-of-the-art results on single-item try-on and leads on multi-item try-on, matching or surpassing leading proprietary systems (Nano Banana Pro, GPT-Image-2, Seedream5 Lite) and open-source models (FLUX.2).

Key Points
  • Reformulates virtual try-on as a multi-reference understanding task instead of traditional mask-based inpainting.
  • Uses a three-stage training pipeline (CPT, SFT, RL) with a hybrid reward model for fine-grained consistency.
  • Outperforms leading proprietary systems (Nano Banana Pro, GPT-Image-2, Seedream5 Lite) and open-source FLUX.2 on both single- and multi-item try-on benchmarks.

Why It Matters

Enables realistic virtual try-on for any fashion item, reducing online shopping returns and transforming e-commerce experiences.

📬 Get the top 10 AI stories daily