Research & Papers

FeDiSyn cuts vision model fine-tuning time by 52.5% with synthetic data

Synthetic images tailored to device distributions slash communication costs by 97.2%.

Deep Dive

Federated fine-tuning (FedFT) of large vision models (LVMs) on distributed, privacy-sensitive devices faces three major hurdles: resource constraints, system heterogeneity, and non-IID data. Existing approaches treat pre-training and fine-tuning separately, missing their inherent connection, and fail to handle weak devices that may still provide critical insights. To address this, researchers from multiple institutions introduce FeDiSyn, a holistic framework that optimizes the entire pipeline from synthetic image pre-training to federated fine-tuning.

FeDiSyn introduces three key innovations: (1) a scaling law for FedFT pre-training that determines the optimal number of synthetic images by balancing pre-training benefit against generation cost; (2) diffusion-based synthetic image generation that mimics device-specific feature distributions to tackle non-IID data; and (3) a contribution-aware LoRA configuration and bandwidth allocation algorithm that prioritizes informative devices while accounting for system heterogeneity. In real-world experiments, FeDiSyn cut completion time by over 52.5% and communication cost by over 97.2%, all while maintaining accuracy comparable to state-of-the-art methods. The paper has been accepted at ICPP 2026.

Key Points
  • FeDiSyn uses a scaling law to find the optimal number of synthetic images for pre-training, balancing cost and benefit.
  • Diffusion-based generation creates synthetic images matching each device's feature distribution to mitigate non-IID data.
  • Contribution-aware LoRA and bandwidth allocation reduce communication cost by 97.2% while preserving accuracy.

Why It Matters

Enables efficient, privacy-preserving fine-tuning of vision models on edge devices with drastically lower latency and bandwidth.

📬 Get the top 10 AI stories daily