Audio & Speech

RealDESED benchmark brings real-world sound detection from 652 homes

5,710 real recordings with multi-annotator labels challenge simulated audio AI benchmarks.

Deep Dive

Current sound event detection (SED) systems are typically trained on simulated soundscapes or web-crawled audio, which fail to capture the variability of real homes — different microphones, placements, background noise, and overlapping events. RealDESED fills this gap with 5,710 recordings (15–35 seconds each) collected by 652 participants in their own homes. The dataset labels 15 common domestic sounds (e.g., vacuum, doorbell, dog bark) with temporally precise annotations.

A key innovation is the multi-annotator labeling scheme: every recording is independently annotated by multiple people, and the validation and test sets undergo extra review for high-quality ground truth. Rich metadata — recording device, placement, environmental labels, and textual descriptions — enables analysis of how these factors affect model performance. The researchers establish a strong transformer baseline achieving a macro-averaged PSDS1 score of 0.731 on the test set. They also explore annotation aggregation strategies, post-processing, and long-form inference.

RealDESED is a step toward bridging the gap between research benchmarks and real-world deployment of SED in smart homes, IoT devices, and assistive technologies. The dataset and baseline code are publicly available on Zenodo and GitHub, and the work has been submitted to the DCASE 2026 Workshop.

Key Points
  • 5,710 recordings from 652 participants' homes, each 15–35 seconds long.
  • Multi-annotator labeling with independent annotations and extra review for validation/test sets.
  • Transformer baseline achieves macro-averaged PSDS1 of 0.731; rich metadata enables deeper analysis.

Why It Matters

Bridges the gap between simulated benchmarks and real-world sound event detection for smart home and IoT applications.

📬 Get the top 10 AI stories daily