Audio & Speech

DAIEN-TTS splits speech from noise to generate real-world acoustic scenes

One-shot voice cloning now controls background noise and reverb separately — no clean audio needed...

Deep Dive

Zero-shot text-to-speech (TTS) has achieved near-human naturalness, but most systems still demand clean, studio-quality speaker prompts — and even then, they either strip away the acoustic environment or tangle it up with the speaker's voice, limiting real-world use. A new paper from Ye-Xin Lu and colleagues at USTC, NICT, and other institutions introduces DAIEN-TTS, an environment-aware zero-shot TTS framework that separates speech, background noise, and reverberation into distinct components. Built on the flow-matching-based F5-TTS model, it uses a speech-environment separation module to decompose environmental audio, then injects those components into a Diffusion Transformer for generation. A cross-speaker conditioning strategy suppresses speaker information leakage from the environment branch.

Training uses simulated data created by mixing clean speech with noise and room impulse responses, with fine-tuning on real-world data to bridge the simulated-to-real gap. At inference, a triple classifier-free guidance mechanism lets users independently control the vocal timbre (via a speaker prompt) and the acoustic environment (via an environment prompt), while a signal-to-noise-ratio adaptation strategy aligns the synthesized voice to the prompts. The team reports that DAIEN-TTS produces speech with high naturalness, strong speaker similarity, and faithful reproduction of noise and reverb — while offering controllability that prior environment-aware TTS systems lack. The work is submitted to IEEE Transactions on Audio, Speech, and Language Processing.

Key Points
  • DAIEN-TTS decouples timbre from environment using separate speaker and environment prompts, enabling independent control of noise and reverb.
  • Built on flow-matching F5-TTS with a speech-environment separation module and triple classifier-free guidance for fine-grained acoustic control.
  • Trained on simulated noisy/reverberant audio with cross-speaker conditioning, plus fine-tuning on real-world data to close the sim-to-real gap.

Why It Matters

Creators can now synthesize speech that naturally matches a film scene, podcast room, or game world without manual audio post-processing.

📬 Get the top 10 AI stories daily