Audio & Speech

NABEATs Uses Reference Noise to Boost Audio AI in Noisy Settings

Reference noise input helps AI focus on target sounds, beating unseen noise types.

Deep Dive

Researchers from MERL and other institutions have introduced NABEATs (Noise-Aware BEATs), a self-supervised learning framework that explicitly handles noise in audio representations. Traditional audio SSL models like BEATs excel on clean signals but degrade under noisy conditions because they encode all sounds equally. NABEATs addresses this by training a model to predict clean BEATs representations from noisy inputs, using an additional reference noise signal. This reference allows the model to learn noise-agnostic features and adapt at inference time to specific noise characteristics, making it robust across diverse real-world environments.

Experimental evaluations show NABEATs achieves significant gains on several downstream tasks (e.g., speech recognition, sound event detection) under both matched and mismatched noise conditions. The model also generalizes to unseen noise types, a critical advantage for deployment. The paper has been accepted at IWAENC 2026, a leading audio processing conference. This noise-aware approach could enable more reliable voice assistants, hearing aids, and automated transcription in challenging acoustic settings.

Key Points
  • Estimates clean BEATs representations from noisy audio using an auxiliary reference noise input.
  • Achieves significant performance improvements across multiple downstream tasks under noisy conditions.
  • Generalizes well to unseen noise types, enhancing real-world deployment robustness.

Why It Matters

Noise-aware SSL like NABEATs enables more reliable audio AI in real-world noisy settings.

📬 Get the top 10 AI stories daily