Research & Papers

Researchers propose D2PO to boost diffusion model quality

New D2PO framework cuts training costs while improving image fidelity by 30%

Deep Dive

Researchers from the University of Seoul have introduced **D2PO (Dynamic Direct Preference Optimization)**, a novel framework designed to optimize diffusion model sampling policies. Published on arXiv (arXiv:2607.06609) and accepted at ECCV 2026, D2PO addresses a key limitation in existing diffusion models: low-NFE (number of function evaluation) student samplers often sacrifice high-frequency texture fidelity to mimic high-NFE teacher models, misaligning with perceptual quality.

The team reformulates sampler optimization as a preference-based alignment problem using Direct Preference Optimization (DPO). They model the sampling policy as an energy-based model (EBM), transforming preference comparisons into tractable energy differences. A novel energy formulation derived from pretrained score networks enables preference evaluation in perturbed spaces that jointly capture structural consistency and fine-grained details. Additionally, D2PO introduces dynamic preferences, where preferred samples progressively improve as the sampling policies are learned, replacing static teacher supervision with an iterative, preference-guided refinement process.

Key Points
  • D2PO uses Direct Preference Optimization (DPO) and energy-based models (EBMs) to optimize diffusion sampling policies
  • Dynamic preferences enable self-improving alignment signals, boosting perceptual quality in low-NFE samplers
  • Experiments demonstrate consistent improvements over traditional regression-based schedulers, with gains up to 30% in texture fidelity

Why It Matters

This framework could reduce compute costs for high-quality image generation while improving visual fidelity in diffusion models.

📬 Get the top 10 AI stories daily