Audio & Speech

Voice AI Accuracy Scores Are Wrong — Especially If You Don't Speak English

Voice AI may be worse than tests suggest — and non-English speakers pay the price.

Deep Dive

Speech recognition — the AI that types what you say, powering voice assistants, video captions and call-center transcripts — is normally graded with a simple score called "word error rate," or WER. Think of it as a spelling test: it counts how many words the machine got wrong. A lower score sounds better, and companies love quoting it.

But there's a wrinkle. Numbers can be written two ways: as digits ("2026") or as words ("two thousand twenty-six"). Some newer AI systems spit out digits; older ones write words out. When you grade one against the other, you're not really measuring mistakes — you're measuring formatting. Most testing tools quietly ignore this problem, and outside English they barely handle it at all.

The authors tested this on Polish, a language with complex grammar, using two real speech datasets: European Parliament recordings and Polish parliamentary sessions. When numbers weren't properly matched up, scores shifted by more than 2 percentage points. That's a big deal, because the gap between two competing voice AI systems on popular public leaderboards is often smaller than that. In other words, the number problem can be louder than the actual difference between products.

The takeaway isn't that voice AI is secretly bad. It's that the scorecards are shakier than they look — especially for anyone who doesn't speak English. If you're using transcription for meeting notes, medical records, customer calls or subtitles, a system's advertised accuracy may not match your real experience. And languages like Polish, Finnish or Turkish, where a single word can carry a lot of meaning, are getting judged with tools built mostly with English in mind. Better testing means better products for everyone, not just English speakers.

Key Points
  • Word error rate (a score counting how many words AI gets wrong) shifts by more than 2 points just based on how numbers are written
  • That 2-point swing is often bigger than the gap between competing voice AI systems on public leaderboards
  • Non-English languages like Polish are graded with tools built mainly for English, so their scores are less trustworthy

Why It Matters

Transcription scores you see in ads may overstate real accuracy, especially for non-English speakers using captions, notes or call tools.

📬 Get the top 10 AI stories daily