Audio & Speech

CrowdioSet and PaRIRset bring live music separation to studio-grade quality

Two open datasets with 4,800 noise tracks and 40 venue impulse responses promise live audio separation.

Deep Dive

Music source separation models 'trained on studio recordings alone' fail in live settings because they ignore venue acoustics, speaker response, and audience noise. To close this gap, Enric Gusó and Xavier Serra (Universitat Pompeu Fabra) built two novel datasets. CrowdioSet offers 4,800 ambient noise tracks from Freesound, plus synthetic sing-along vocals generated via zero-shot singing voice conversion from the MUSDB18 and MOISESDB datasets. The authors show that training with CrowdioSet significantly improves denoising of live vocal recordings, yielding superior results in both objective metrics (like SDR) and subjective listening tests.

PaRIRset complements this with a stereo impulse response (IR) dataset captured at 40 professional concert venues using a microphone array. Adding these real-world room impulse responses to the training pipeline boosts MSS performance compared to using only speech-enhancement RIRs. The combined approach helps models generalize to the acoustic chaos of live shows—crowd clapping, echo, and PA system coloration. The researchers have released the datasets, model weights, and code openly, making it easy for developers to fine-tune existing separation systems. Accepted to ISMIR26, this work is a practical step toward reliable live audio processing for remixing, karaoke, and live-streaming tools.

Key Points
  • CrowdioSet adds 4,800 real ambience tracks from Freesound plus synthetic sing-alongs created via zero-shot singing voice conversion from MUSDB18 and MOISESDB.
  • PaRIRset captures stereo impulse responses from 40 professional concert venues using a microphone array, improving MSS model generalization.
  • The method outperforms studio-only and speech-RIR baselines for live music source separation, with code, weights, and datasets released openly.

Why It Matters

Live concert remixing, karaoke apps, and broadcast audio need separation that survives crowds and reverb — these datasets make it possible.

📬 Get the top 10 AI stories daily