New framework transcodes any spatial audio format in real time
A universal transcoder turns Ambisonics and raw mic arrays into any playback system.
A new paper from Archontis Politis, Janani Fernandez, and Leo McCormack, submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing, presents a unified framework for arbitrary spatial audio transcoding. The system takes either Ambisonic signals or raw microphone array captures and estimates time-frequency spatial metadata describing primary source components and an ambience component with its own angular power distribution. This metadata is used to construct spatial covariances for any target playback format (e.g., headphones, speaker arrays) and derive optimal mixing matrices. The framework also supports independent rotations of both capture and playback setups, enabling flexible alignment in VR or mixed-reality applications.
Real-time implementations were compared to state-of-the-art parametric renderers in listening tests using simulated scenes from Ambisonic, spherical, and head-worn arrays. Results showed perceptual advantages for the proposed method across diverse content and receiver configurations, particularly for lower-order Ambisonics and geometrically constrained microphone arrays. This transcoding approach eliminates the need for format-specific solutions, potentially standardising how spatial audio is captured, stored, and reproduced in professional and consumer systems.
- Estimates time-frequency spatial metadata from Ambisonic or raw microphone array captures.
- Handles independent rotation of both capture and playback coordinate systems.
- Outperforms existing parametric renderers in listening tests for lower-order and constrained arrays.
Why It Matters
Simplifies spatial audio production and playback across any device, from VR headsets to home theatres.