dots.tts.edit enables surgical precision in AI voice editing
New model lets you edit speech with XML-style tags for surgical precision.
A team led by Hankun Wang from Shanghai Jiao Tong University has developed dots.tts.edit, a groundbreaking speech editing model that delivers surgical precision through transcript-grounded structural edit instructions. Unlike ambiguous natural language prompts, this approach uses XML-style tags to explicitly define edit operations and their target regions in the transcript, eliminating the need for explicit timestamp alignment.
The model, built on a continuous autoregressive foundation, supports four key speech-creation controls: lexical content (text editing), affective expression (emotion editing), prosody (pitch and speaking-rate delivery), and temporal phrasing (pause editing). The researchers introduced doteBench, a bilingual evaluation suite that measures instruction following, local preservation, and audio quality across these controls and their combinations. Experiments show leading performance in instruction adherence and preservation while maintaining audio quality comparable to existing open-source systems.
- dots.tts.edit uses XML-style tags for explicit, unambiguous speech editing instructions instead of natural language prompts
- Supports four control types: lexical content, affective expression, prosody, and temporal phrasing
- Maintains audio quality comparable to existing open-source systems while achieving leading instruction following and preservation metrics
Why It Matters
Revolutionizes voice content creation with surgical precision editing, enabling professionals to fine-tune audio with unprecedented control and reliability.