Brazilian researchers slash audio AI compute costs 14x
New AEROMamba architecture cuts audio enhancement costs by 20x with psychoacoustic training
Researchers Wallace Abreu, Bernardo V. Miranda, and Luiz W. P. Biscainho from Brazilian institutions have developed AEROMambaP, a novel audio enhancement architecture that achieves dramatic computational efficiency while improving sound quality. Published in the Journal of the Audio Engineering Society (v74, n6), their work introduces Mamba state-space models to replace traditional attention mechanisms and LSTM layers in audio super-resolution tasks.
The key innovation lies in incorporating differentiable psychoacoustic loss derived from the Perceptual Audio Quality Measure (PAQM). During training, AEROMambaP requires 2-4x less GPU memory than the baseline AERO architecture. During inference, it delivers a 14x speedup while consuming only one-fifth of the GPU memory. In subjective listening tests, the model achieved 15% higher perceived quality scores when upscaling piano datasets and MUSDB18 from 11.025 kHz to 44.1 kHz. For compressed audio restoration, the variant AEROMambaPS showed 52% higher quality ratings when enhancing 32 kbps MP3 files.
- AEROMambaP achieves 14x faster inference with 80% lower GPU memory usage compared to baseline AERO
- Includes differentiable PAQM loss for perceptual audio quality optimization
- Achieves 15% higher perceived quality in audio upscaling and 52% better restoration of compressed MP3 audio
Why It Matters
Audio AI just got 20x cheaper to run while delivering higher quality results for both lossless and compressed audio scenarios