Audio & Speech

New probing framework reveals disentanglement flaws in acoustic teleportation codecs

Acoustic teleportation codecs leak room acoustics into speech embeddings, study finds.

Deep Dive

A new paper from Philipp Grundhuber and Emanuël A. P. Habets (FAU Erlangen-Nuremberg) tackles a fundamental challenge in neural audio codecs: how to reliably evaluate whether latent representations truly disentangle speech content, speaker identity, and acoustic environment. Current methods rely on cross-reconstruction quality—feeding the wrong combination of latents into a decoder—but this can miss subtle leakage between partitions. The authors extend a probing-based framework that directly measures how much information about room acoustics (reverberation time, clarity, direct-to-reverberant ratio) and speaker identity can be predicted from each latent subspace. The gap between intended and unintended partitions serves as the disentanglement metric.

Applied to an acoustic teleportation codec, the probing reveals speaker identity is largely confined to its designated partition, but room acoustics leak significantly into the speech content embeddings due to the model's training objective. Strikingly, the acoustic embeddings themselves predict room parameters within 0.02 seconds of supervised baselines, indicating that physically meaningful structure emerges without explicit labeling. This work (accepted at Interspeech 2026) provides a more rigorous evaluation toolkit for voice conversion and acoustic teleportation systems, highlighting that cross-reconstruction alone can mask critical information leakage that degrades real-world performance.

Key Points
  • Probing-based evaluation regresses room-acoustic parameters (RT60, C50, DRR) and classifies speaker identity to measure disentanglement quality.
  • Acoustic embeddings estimate room parameters within 0.02s of supervised baselines, proving unsupervised learning of physically meaningful structure.
  • Current cross-reconstruction metrics miss leakage—here, room acoustics leak into speech embeddings while speaker identity stays confined.

Why It Matters

Improves acoustic teleportation and voice conversion by identifying latent leakage that cross-reconstruction can't detect.

📬 Get the top 10 AI stories daily