Real-time MRI speech enhancement study reveals surprising trade-offs
Three leading speech enhancement systems for real-time MRI evaluated across 15 scenarios with mixed results
A new arXiv preprint (arXiv:2608.16125) from researchers at USC, UT Dallas and others examines how three off-the-shelf speech enhancement systems perform on real-time MRI (rtMRI) audio contaminated by scanner noise. The team evaluated Denoiser, PASE, and RE-USE across five rtMRI corpora using multiple evaluation metrics including ASR performance, intelligibility measures, and speaker-preservation metrics.
The study found endpoint-dependent results where higher predicted-quality scores didn't correlate with better ASR performance. In 15 corpus-recognizer comparisons, RE-USE yielded lower word-error-rate estimates in 11 cases, while Denoiser increased error rates in 13. PASE and RE-USE showed improvements in phone agreement and intelligibility in paired-noise tests, but Denoiser reduced speaker-embedding similarity despite improving some objective metrics. The researchers conclude that enhanced rtMRI audio should be treated as a task-specific transformed derivative rather than a universal improvement.
- Three speech enhancement systems (Denoiser, PASE, RE-USE) tested across 15 rtMRI corpus-recognizer combinations
- RE-USE improved ASR in 11/15 cases while Denoiser degraded performance in 13/15 scenarios
- No system was universally superior across all evaluation metrics and corpora
Why It Matters
Critical insight for medical AI and speech processing teams: speech enhancement isn't one-size-fits-all for rtMRI applications