Deep learning ReTM estimation beats covariance method for noisy speech
Three neural frameworks estimate Relative Transfer Matrix for multi-source, multi-mic audio with higher accuracy.
Researchers Yalegama and Manamperi proposed deep learning-based frameworks for estimating the Relative Transfer Matrix (ReTM), a recently introduced generalization of the relative transfer function for multiple receivers and sources. Their three supervised approaches—time-domain and short-time frequency transform domain convolutional networks, plus an LSTM-based recurrent neural network—achieved more accurate ReTM estimation than the covariance-based method across five objective metrics. The frameworks also showed speech enhancement performance on par with the baseline method. The paper was accepted to Interspeech 2026.
- Three deep learning frameworks proposed: time-domain CNN, STFT-domain CNN, and LSTM-based RNN for ReTM estimation
- Outperformed covariance-based method across five objective metrics for multi-source, multi-microphone scenarios
- Speech enhancement performance matches baseline method; accepted at Interspeech 2026 and available on arXiv (2608.11627)
Why It Matters
Neural ReTM estimation could make hearing aids and smart speakers far more robust in noisy, multi-speaker environments.