Audio & Speech

Veler & Gannot's unfolded REM network tracks speakers with lower RMSE

New deep learning approach adaptively updates tracking policy in mild reverberant environments.

Deep Dive

Tracking a moving speaker in reverberant spaces is a classic challenge in audio processing. Traditional recursive expectation-maximization (REM) algorithms rely on fixed step-size decay schedules, which limit adaptability in dynamic acoustic conditions. Now, researchers Rina Veler and Sharon Gannot from Bar-Ilan University have introduced a deep unfolded REM network that learns an adaptive update policy. Their method 'unfolds' the iterative REM procedure into differentiable layers, allowing the network to optimize tracking as part of an end-to-end training framework.

The key innovation is a Step Size Network that leverages Feature-wise Linear Modulation (FiLM) and Positional Encoding (PE) to dynamically adjust recursion weights based on temporal context and convergence state. In single-speaker tracking experiments under mild reverberation, the proposed network outperforms the classical CREM baseline, achieving a lower RMSE (root mean square error) without requiring a spatial grid search for position mapping. This work, presented at IWAENC 2026, demonstrates the potential of learned optimization for real-time acoustic tracking in smart rooms, hearing aids, and autonomous systems.

Key Points
  • Unfolds iterative REM algorithm into differentiable layers, enabling end-to-end learning of update policy.
  • Introduces Step Size Network using FiLM and PE to dynamically adjust recursion weights based on temporal context.
  • Achieves lower RMSE than classical CREM baseline in single-speaker tracking under mild reverberation.

Why It Matters

Enables more accurate, adaptive speaker tracking for smart audio systems without manual parameter tuning.

📬 Get the top 10 AI stories daily