Open Source

ServiceNow benchmarks ElevenLabs, Gemini, AssemblyAI on bilingual code-switched speech

Top ASR models tested on Spanish-English, French-English, German-English code-switching for enterprise use.

Deep Dive

ServiceNow AI researchers have released a new benchmark addressing a critical gap in voice agent capabilities: handling code-switched speech common among bilingual customers. Their study, published June 9, 2026, focuses on automatic speech recognition (ASR) for enterprise settings, where transcription errors cascade into misrouted tickets or misunderstood policies. They created a dataset of 918 utterances spanning four language pairs—Spanish-English, French-English, Canadian French-English, and German-English—covering realistic HR and IT support scenarios (e.g., benefits inquiries, password resets). Utterances were generated from parallel corpora using GPT-5 for code-switching, verbalized via LLM, and synthesized with ElevenLabs Multilingual V2, then verified by native-speaking linguists. The team evaluated seven ASR systems including ElevenLabs Scribe V2, Gemini 3 Flash, and Assembly AI Universal 3-Pro using three metrics: Word Error Rate (WER), Semantic WER (measuring meaning preservation), and Answer Error Rate (AER) for downstream task accuracy.

The results show that code-switching costs vary by language pair and model. ElevenLabs Scribe V2, Gemini 3 Flash, and Assembly AI Universal 3-Pro consistently outperformed others across all metrics. Spanish-English code-switching proved generally easier than German-English for most models. The benchmark, released via AU-Harness, provides a standardized way to evaluate voice agents on bilingual speech—a growing need as over half the world's population speaks multiple languages. For enterprises serving diverse customer bases, these findings highlight that not all ASR systems handle mixed-language conversations equally, and that ignoring code-switching can lead to significant operational failures in automated support pipelines.

Key Points
  • Seven ASR systems benchmarked on 918 code-switched utterances across four language pairs (Spanish-English, French-English, Canadian French-English, German-English).
  • Top performers: ElevenLabs Scribe V2, Gemini 3 Flash, Assembly AI Universal 3-Pro; cost of code-switching varies by language pair.
  • Metrics include Word Error Rate, Semantic WER (meaning preservation), and Answer Error Rate for downstream task accuracy in HR/ITSM scenarios.

Why It Matters

Enterprises serving bilingual customers can now identify which ASR models handle mixed-language speech reliably, reducing operational errors.

📬 Get the top 10 AI stories daily