Viral Wire

Alibaba's Qwen Audio 3.0 beats OpenAI in speech-to-speech benchmark

New model scores 84.1%, surpassing OpenAI's 79.1% in real-time speech AI test

Deep Dive

Alibaba Cloud released its Qwen-Audio-3.0-Realtime model family on July 28, 2026, targeting real-time voice interaction. The flagship Qwen-Audio-3.0-Realtime Plus achieved an overall score of 84.1% on the Artificial Analysis Speech-to-Speech Index, topping OpenAI's GPT-Realtime-2.1 High at 79.1%. This benchmark measures end-to-end voice conversation quality, including latency, accuracy, and naturalness.

The Qwen lineup includes multiple tiers and introduces full-duplex communication — both parties can speak and listen simultaneously without turn-taking. It also features function calling for integrating with external APIs and custom cloned voices for personalization. Alibaba positions these models as direct competitors to OpenAI's GPT-Realtime series, offering a compelling alternative for enterprises building voice assistants, customer service bots, and real-time translation tools.

Key Points
  • Launched on July 28, 2026, with multiple tiers including Qwen-Audio-3.0-Realtime Plus
  • Scored 84.1% on Artificial Analysis Speech-to-Speech Index vs OpenAI's 79.1%
  • Supports full-duplex, function calling, and custom cloned voices for enterprise use

Why It Matters

Real-time voice AI competition heats up, giving enterprises a strong alternative to OpenAI for live conversations.

📬 Get the top 10 AI stories daily