Viral Wire

Microsoft's MAI-Transcribe-1.5 delivers 5x faster transcription with 30% fewer errors

Transcribes an hour of audio in under 15 seconds across 43 languages.

Deep Dive

Microsoft AI announced MAI-Transcribe-1.5, an automatic speech recognition (ASR) model built in-house. It expands language coverage from 25 to 43 languages, including new South Asian and European languages, without accuracy trade-offs. On the FLEURS benchmark, Microsoft claims best-in-class WER across all 43 languages; on the Artificial Analysis leaderboard, it posts a 2.4% WER (ranked #3). The model transcribes an hour of audio in under 15 seconds, achieving up to 5x faster inference than Gemini 3.1, Scribe v2, and GPT-4o-Transcribe, and up to 5.7x faster than its predecessor for long-form audio.

A standout feature is keyword (entity) biasing. Users can supply up to 200 domain-specific keywords (e.g., product names, medical terms, or internal acronyms), and the model intelligently biases its predictions using context—reducing WER by 30% on FLEURS. For example, it correctly recovers “Shaun,” “Aoife,” and “Xochitl” when biased, versus generic transcriptions. MAI-Transcribe-1.5 also features automatic language identification and is integrated into Microsoft Copilot, Teams, GitHub, and Dynamics 365 Contact Centre, with availability on Azure Foundry. Use cases include video captions, meeting transcription, call center analytics, and voice agent pipelines.

Key Points
  • Supports 43 languages with a single model and achieves a 2.4% WER on the Artificial Analysis leaderboard.
  • Transcribes one hour of audio in under 15 seconds—up to 5.7x faster than previous generation and 5x faster than competitors.
  • Keyword biasing with up to 200 custom terms reduces WER by 30% on FLEURS, enabling accurate transcription of domain-specific vocabulary.

Why It Matters

Enterprise transcription just got faster, more multilingual, and customizable for domain-specific accuracy at scale.

📬 Get the top 10 AI stories daily