Tencent Cloud and Inworld AI merge TTS with global RTC for lifelike voice AI
Inworld's #1 ranked TTS now integrates with Tencent's 3,200-node global network for sub-130ms voice.
Tencent Cloud and Inworld AI today announced a strategic partnership that integrates Inworld's state-of-the-art text-to-speech (TTS) models with Tencent's enterprise-grade Real-Time Communication (RTC) infrastructure. Inworld's TTS, ranked #1 on the Artificial Analysis Speech Arena, delivers ultra-realistic, context-aware speech with emotional nuance and precise voice cloning. Key specs include sub-130ms first-chunk latency, support for over 100 languages, and instant cross-lingual conversion while preserving a consistent speaker voice. Developers can clone voices, design custom voices, and control vocal style directly through the Tencent RTC console and SDK.
Tencent RTC provides the global backbone with more than 3,200 nodes and sub-300ms worldwide latency, complemented by AI noise suppression and weak-network resilience. This combination ensures that conversational AI applications run seamlessly even in challenging connectivity environments. The partnership also launches a Conversational AI Demo featuring recommended voices for different languages and use cases. According to Wison Xie, Head of Tencent RTC Product, production-grade voice AI requires both a capable model and a capable network — and this partnership delivers a one-stop path from prototype to production.
- Inworld TTS ranked #1 on Artificial Analysis Speech Arena for expressive, real-time voice.
- Sub-130ms first-chunk latency, 100+ languages, cross-lingual voice cloning.
- Tencent RTC provides 3,200+ global nodes with sub-300ms latency and AI noise suppression.
Why It Matters
Developers can now deploy emotionally intelligent, real-time voice agents globally with minimal infrastructure overhead.