Rec-RIR: New AI identifies room acoustics from speech alone
No reference mic needed—deep network extracts room impulse response from noisy recordings.
Researchers introduced Rec-RIR, a multi-task deep neural network that blindly identifies room impulse responses (RIRs) from reverberant speech recordings. Using convolutive transfer function (CTF) approximation, the model sequentially removes noise and reverberation, then reconstructs the CTF filter. A pseudo-intrusive measurement step converts this into standard RIR. The method achieves state-of-the-art performance (accepted at Interspeech 2026).
- Rec-RIR uses a multi-task neural net to jointly denoise, dereverberate, and estimate room acoustics.
- It converts estimated convolutive transfer functions (CTF) into standard room impulse responses via a pseudo-intrusive measurement step.
- Accepted at Interspeech 2026; achieves state-of-the-art results in blind RIR identification.
Why It Matters
Real-time, device-free room acoustics estimation unlocks better voice interfaces and immersive audio in any space.