Audio & Speech

MSA-EchoLite slashes AI echo cancellation compute by 50%

New MSA-EchoLite model uses 0.2M parameters to achieve 99.1% PESQ with 100M FLOPs/s...

Deep Dive

A team of researchers from Nanjing University and the University of Passau has unveiled MSA-EchoLite, a breakthrough lightweight acoustic echo cancellation (AEC) framework that dramatically reduces computational requirements while maintaining high performance. Published on arXiv as arXiv:2608.03650, the work introduces an asymmetric dual-branch encoder architecture with an echo-aware frequency-time modulation (EAM) module that enriches compressed bottleneck representations by modeling discrepancy and correlation cues between microphone and echo-related features.

The innovation delivers exceptional efficiency: MSA-EchoLite achieves 99.1% of the PESQ score of heavier frequency-domain counterparts while requiring only 100M FLOPs per second and just 0.2 million parameters. With just 26.1% additional FLOPs over its non-EAM variant, the EAM-enhanced version not only matches but surpasses traditional frequency-domain models in SDR metrics. This represents a significant advancement for real-time AEC applications where computational resources are constrained, from video conferencing systems to smart speakers and AR/VR devices.

Key Points
  • MSA-EchoLite achieves 99.1% PESQ of frequency-domain models using only 100M FLOPs/s and 0.2M parameters
  • The echo-aware frequency-time modulation module adds just 26.1% FLOPs while improving performance
  • Outperforms state-of-the-art lightweight AEC models in both PESQ and SDR metrics

Why It Matters

Enables real-time high-quality echo cancellation on edge devices with minimal compute overhead.

📬 Get the top 10 AI stories daily