VoiceDesigner uses diffusion to edit voices like photos
New unified model generates and edits voices with 20% better prompt alignment
A team led by Jiarui Hai at Johns Hopkins University and including researchers from Carnegie Mellon, Adobe, and other institutions introduced VoiceDesigner, a unified framework for text-to-voice generation and editing. Published on arXiv (arXiv:2608.13613), the paper addresses two key challenges in current TTV systems: limited voice diversity and inflexible editing capabilities.
VoiceDesigner tackles these issues with a hybrid data pipeline combining digital signal processing and speech generation models to create a diverse dataset covering real-world and fictional voices. At its core, the system uses a diffusion transformer with architectural improvements to better handle complex conditioning and enable unified voice generation and editing tasks. The model achieves superior prompt alignment with both voice descriptions and editing instructions while maintaining competitive perceptual quality and voice usability, outperforming state-of-the-art TTV models in evaluations.
- Unified framework combines text-to-voice generation and editing in a single model
- Hybrid data pipeline creates diverse voices (real-world and fictional) with DSP and speech models
- Diffusion transformer architecture improves conditioning and multi-task performance
Why It Matters
Enables creators to generate and fine-tune voices with photographic precision, revolutionizing audio production workflows