Mesh2HRTF synthetic HRTFs match measured spatial audio across 200 subjects
Synthetic HRTFs from Mesh2HRTF rival measured ones in VR localization, with no torso geometry needed.
A new arXiv preprint by Katarina Poole and Lorenzo Picinali tackles a long-standing bottleneck in personalized spatial audio: measuring Head-Related Transfer Functions (HRTFs) for every listener. Using the boundary element method (BEM) via the Mesh2HRTF pipeline, the team generated synthetic HRTFs from head-and-torso mesh models and compared them against individually measured HRTFs and the generic KEMAR manikin across 200 subjects from the Extended SONICOM dataset. Numerically, synthetic HRTFs achieved lower interaural time and level differences (ITD/ILD) errors than KEMAR, though residual spectral distortion concentrated at low rear elevations—an artifact of omitting torso geometry from the synthesis pipeline. Two computational localization models mirrored this pattern, predicting errors intermediate between measured and KEMAR.
Crucially, behavioral validation overturned these numerical predictions. In a virtual reality localization task with 20 participants, synthetic HRTFs matched measured HRTFs on every polar coordinate metric, while KEMAR was significantly worse. Interestingly, behavioral errors clustered around the front-back midline regardless of condition, not at the low elevations flagged by the numerical models. A separate spatial release from masking test with 18 participants found no significant effect of HRTF type. The authors conclude that high-resolution synthetic HRTFs can preserve behavioral localization performance, suggesting that numerical discrepancies don't necessarily translate to perceptible degradation. This work (arXiv:2608.16722, submitted to JASA) could accelerate scalable, personalized spatial audio for VR, gaming, and telepresence.
- Mesh2HRTF synthetic HRTFs beat KEMAR on interaural time/level differences across 200 subjects.
- VR localization task (N=20) showed synthetic HRTFs match measured on every polar metric; KEMAR significantly worse.
- Missing torso geometry causes low-elevation spectral errors, but behavioral results show no real-world localization penalty.
Why It Matters
Scalable synthetic HRTFs could eliminate costly per-user measurements, enabling mass-market personalized spatial audio for VR, AR, and gaming.