Audio & Speech

Two-stage AVSE model boosts real-world speech enhancement with CLIP matching

Audio-only separation plus CLIP face matching beats real-world cocktail party problem

Deep Dive

Audio-visual speech enhancement (AVSE) aims to isolate a target speaker from multi-speaker mixtures using visual cues. While simulated datasets show strong results, real-world performance often collapses when speaker overlap, reverberation, and visual degradation co-exist. To address this, Tongtao Ling and Zhong-Qiu Wang submitted a "separate first, then associate" approach to the Real-World AVSE Challenge at ISCSLP 2026. Their method decouples the problem into two stages, bypassing the fragile end-to-end joint training commonly used in prior work.

The first stage uses a trained audio-only model (no visual inputs) to separate the mixture into individual speaker signals. The second stage employs an audio-visual CLIP model to compute cross-modal similarity between each separated speech signal and the target speaker's facial video, selecting the one with the highest match. This modular design lets each stage be optimized independently and leverages strong pre-trained models. Evaluation on the challenge dataset confirms the approach's effectiveness, offering a practical solution for real-world AVSE under naturally co-existing acoustic and visual distortions.

Key Points
  • Two-stage design: audio-only separation first, then audio-visual CLIP-based association
  • No visual cues used in the separation stage, reducing reliance on degraded video
  • Validated on the Real-World AVSE Challenge (ISCSLP 2026) dataset with real-world conditions

Why It Matters

This decoupled approach makes AVSE practical for hearing aids, video calls, and speech recognition in noisy, reverberant real environments.

📬 Get the top 10 AI stories daily