MSA-EchoLite slashes AI echo cancellation compute by 50%
New MSA-EchoLite model uses 0.2M parameters to achieve 99.1% PESQ with 100M FLOPs/s...
A team of researchers from Nanjing University and the University of Passau has unveiled MSA-EchoLite, a breakthrough lightweight acoustic echo cancellation (AEC) framework that dramatically reduces computational requirements while maintaining high performance. Published on arXiv as arXiv:2608.03650, the work introduces an asymmetric dual-branch encoder architecture with an echo-aware frequency-time modulation (EAM) module that enriches compressed bottleneck representations by modeling discrepancy and correlation cues between microphone and echo-related features.
The innovation delivers exceptional efficiency: MSA-EchoLite achieves 99.1% of the PESQ score of heavier frequency-domain counterparts while requiring only 100M FLOPs per second and just 0.2 million parameters. With just 26.1% additional FLOPs over its non-EAM variant, the EAM-enhanced version not only matches but surpasses traditional frequency-domain models in SDR metrics. This represents a significant advancement for real-time AEC applications where computational resources are constrained, from video conferencing systems to smart speakers and AR/VR devices.
- MSA-EchoLite achieves 99.1% PESQ of frequency-domain models using only 100M FLOPs/s and 0.2M parameters
- The echo-aware frequency-time modulation module adds just 26.1% FLOPs while improving performance
- Outperforms state-of-the-art lightweight AEC models in both PESQ and SDR metrics
Why It Matters
Enables real-time high-quality echo cancellation on edge devices with minimal compute overhead.