Audio & Speech

MERL wins Real-TSE Challenge by prioritizing data over architecture

A four-stage training pipeline and clever data cleaning beat top competitors.

Deep Dive

MERL's submission to the Real-TSE Challenge proves that data quality can trump architectural novelty. Instead of designing a new model, the team built upon the baseline and concentrated on a four-stage training strategy. Stage one pre-trained on synthetic fully overlapped speech mixtures. Stage two introduced simulated multi-talker conversations with added noise and reverberation—applied to both mixtures and enrollment utterances. Stage three adapted the model to real-world far-field noisy recordings using pseudo-targets derived from processed close-talk microphone signals. The final stage fine-tuned for the challenge's specific track conditions.

This practical approach earned MERL first place in Track 2, demonstrating that careful data preparation is critical for target speech extraction in real-world conditions. The team also investigated metric robustness, finding that DNSMOS (speech quality) and speaker similarity scores can be driven to extreme values without improving actual performance measures like token error rate or VAD-based F1 score. This warns against over-reliance on such metrics and highlights the need for more robust evaluation methods.

Key Points
  • MERL won Track 2 of the Real-TSE Challenge without a new model architecture, solely by improving data preparation and training pipeline.
  • The system was trained in four stages: synthetic mixtures, simulated conversations with noise/reverb, real far-field recordings with pseudo-targets, and final fine-tuning.
  • Research showed that DNSMOS and speaker similarity metrics can be easily over-optimized (driven to extreme values) without improving actual token error rate or VAD F1 score.

Why It Matters

Proves that data quality and training strategy can beat model innovation in speech separation—a key lesson for deployable AI systems.

📬 Get the top 10 AI stories daily