New physics-informed diffusion model boosts spatial audio fidelity
DiffM2A delivers 20% better spatial audio from sparse mic arrays...
DiffM2A, a geometry-adaptive conditional diffusion framework introduced by Xiang Zhou, Zhengqiao Zhao, Zhengding Luo, and Wen Zhang, targets a persistent problem in spatial audio: higher-order Ambisonic encoding for sparse microphone arrays. Because these arrays are often irregular and shaped by device-specific boundary conditions, spherical-harmonic encoding becomes ill-conditioned — inverse filtering amplifies noise, while deterministic neural encoders risk overfitting to array-specific responses. DiffM2A's Geometry-Adaptive Spherical Harmonic Projection (GASHP) front-end builds boundary-aware steering functions with an energy-normalized modal projection, sidestepping explicit pseudo-inverse computation. A dual-branch Elucidated Diffusion Model then estimates complex Ambisonic coefficients using both raw microphone spectra and GASHP features, while sound intensity and rotational equivariance losses sharpen inter-channel phase consistency. Tested on first- and second-order Ambisonic encoding tasks with simulated room acoustics and real-world LOCATA recordings, DiffM2A outperformed
- DiffM2A uses physics-informed diffusion to boost Ambisonic encoding quality by 20% over conventional methods
- The Geometry-Adaptive Spherical Harmonic Projection (GASHP) front-end eliminates pseudo-inverse computation for sparse mic arrays
- Model maintains spatial coherence across variable microphone topologies and unseen layouts
Why It Matters
Enables high-fidelity spatial audio in compact devices like AR/VR headsets and hearables where mic arrays are sparse and irregular.