Audio & Speech

NTT's C2D Projection Outperforms GSS on Real-World Speech Enhancement

New method uses real microphone pairs instead of simulations to train cleaner speech models.

Deep Dive

Training neural networks for speech enhancement in distant-recording scenarios typically requires paired distorted and clean reference signals. Simulated data often mismatches real conditions, limiting accuracy. To solve this, NTT's Tomohiro Nakatani and colleagues introduce Close-to-Distant microphone Projection (C2D projection). The method leverages real recordings from both close and distant microphones, estimating an optimal projection matrix via a variant of the Parametric Multichannel Wiener Filter (PMWF). This matrix transforms close-microphone inputs into clean targets that align with the distant microphone's perspective, while simultaneously performing denoising.

Experiments on the challenging CHiME6 dinner party ASR dataset show that a neural network trained with C2D-projected data outperforms Guided Source Separation (GSS), the current state-of-the-art, when using GSS's enhanced output as an auxiliary input under oracle diarization. The work demonstrates that generating training targets from real recordings can bridge the simulation-to-reality gap, offering a practical path to more robust speech enhancement in noisy environments like conferences, smart homes, and hands-free devices.

Key Points
  • C2D projection generates training pairs from real close and distant microphone recordings, avoiding simulation mismatches.
  • The method uses a variant of the Parametric Multichannel Wiener Filter (PMWF) to estimate a projection matrix that aligns and denoises signals.
  • On the CHiME6 dinner party ASR task, a neural net trained with C2D data outperformed the previous state-of-the-art (GSS) with oracle diarization.

Why It Matters

Enables more accurate speech enhancement in real noisy environments like conferences, dinner parties, and smart home devices.

📬 Get the top 10 AI stories daily