Research & Papers

Transsion's New AI Transcribes Multilingual Meetings and Knows Who Said What

Meeting notes that label every speaker — in several languages at once.

Deep Dive

The Transsion Speech Team submitted a system to Task 1 of the MLC-SLM 2026 Challenge, which focuses on speaker-attributed transcription for multilingual conversational speech. Their framework cascades three components: a speaker diarization module built on DiariZen, which produces speaker-homogeneous segments through local speaker activity estimation and global speaker clustering; a long-form multilingual ASR module based on Qwen3-Omni; and a speaker-transcription fusion module. An external CTC-based alignment model supplies precise word- and character-level timestamps, and the fusion module combines diarization outputs with those timestamped transcriptions to generate speaker-attributed STM outputs. On the official evaluation set, the submitted system achieved a tcpMER of 15.41% and ranked second among all participating teams.

Key Points
  • It's speech-to-text that also labels speakers, so a transcript reads like a script with names instead of a wall of anonymous words.
  • It handles multiple languages in one conversation and placed second in an international competition judged on real recordings.
  • The system still gets about one in seven words or speaker labels wrong, and it isn't available as a product yet.

Why It Matters

Could mean automatic, labeled meeting notes and searchable call recordings in the languages you actually speak.

📬 Get the top 10 AI stories daily