Open Source

RedNote's dots.tts 2B: Open-Source TTS with Zero-Shot Voice Cloning

A 2B-parameter model that clones voices from just a few seconds of audio.

Deep Dive

RedNote (Xiaohongshu) has open-sourced dots.tts 2B, a state-of-the-art text-to-speech model with 2 billion parameters. Its fully continuous architecture eschews traditional codec tokens, allowing for smoother and more natural speech synthesis at 48 kHz. The model is released under the permissive Apache 2.0 license, with code, a demo page, and a technical report available on GitHub and arXiv. This design eliminates the need for discrete codec layers, preserving finer acoustic details and enabling better prosody control.

A standout feature is zero-shot voice cloning, which allows dots.tts to replicate a speaker's voice from just a few seconds of reference audio without any fine-tuning. The model also skips the conventional phoneme pipeline, directly mapping text to speech waveforms. This reduces engineering complexity and latency, making it practical for real-time applications. With its high sampling rate, creative developers can use dots.tts for audiobook generation, virtual assistants, and personalized voice interfaces. The open-source release invites community contributions and further research into continuous speech representations.

Key Points
  • 2 billion parameter TTS model released under Apache 2.0 license
  • Fully continuous architecture — no codec tokens — for higher fidelity 48kHz synthesis
  • Zero-shot voice cloning with direct text-to-speech, no phoneme preprocessing needed

Why It Matters

Pushes open-source TTS forward with high-fidelity voice cloning, enabling developers to build realistic speech applications.

📬 Get the top 10 AI stories daily