Audio & Speech

VeRe-Flow: Noise-Robust Bandwidth Expansion via Clean Guidance

New flow matching method achieves lowest distortion and highest speech quality scores

Deep Dive

VeRe-Flow addresses a fundamental challenge in audio processing: reconstructing high-fidelity wideband speech from noisy low-resolution inputs. Traditional flow matching struggles under noise because velocity estimation becomes ambiguous. The KAIST team introduces two novel clean-guidance mechanisms: velocity contrastive regularization attracts predicted velocity toward clean trajectories while repelling noisy ones, and representation alignment forces intermediate features to match clean self-supervised learning representations. This dual-level supervision enables the generative process to bypass noise contamination and produce clean wideband speech.

Experimental results show VeRe-Flow outperforms all baselines on both objective and subjective metrics. It achieves the lowest Log-Spectral Distortion (LSD) and highest DNSMOS OVRL (Deep Noise Suppression Mean Opinion Score Overall) among all compared methods, and the highest MOS among generative baselines. The paper has been accepted to Interspeech 2026, confirming the work's significance. Practical applications include real-time speech enhancement for voice assistants, hearing aids, and communications systems where degraded audio must be restored without artifacts.

Key Points
  • Velocity contrastive regularization steers flow matching toward clean speech trajectories while repelling noisy ones
  • Representation alignment matches intermediate features with clean self-supervised learning embeddings
  • Achieves best LSD and DNSMOS OVRL scores, with highest MOS among generative models, accepted to Interspeech 2026

Why It Matters

Enables clean wideband speech reconstruction from noisy low-res inputs, improving voice assistants, hearing aids, and communications

📬 Get the top 10 AI stories daily