Audio & Speech

New AI Transcribes Meetings Live and Knows Who Said What

Imagine live meeting notes that always know who's talking—even with interruptions.

Deep Dive

If you've ever struggled to follow a video call or wanted accurate live captions, this AI news is for you. Researchers have unveiled VibeVoice-ASR-Streaming, an AI system that can transcribe conversations in real time and also say which person is speaking. In tech terms, it merges speech-to-text with speaker recognition into one fluid process. That means software like meeting notebooks, voice assistants, or phone call transcripts could soon label each line of text with the right name as the words are spoken, not moments later.

What makes this different from what existed before? Older systems treated transcription and speaker detection as separate steps. They would first turn the whole audio clip into text, then process it again to figure out who said what. That two-step approach works fine for recorded files, but it's too slow for live assistants and real-time tools. The new model works on the fly, using short chunks of audio and a little bit of future sound to keep up with the conversation, like reading a text as it's being typed rather than after the whole message arrives.

The creators trained two versions of the model: a larger one with 7 billion parameters (roughly the number of "digital connections" the AI uses to make decisions) and a smaller 1.5 billion version. On test sets of five different audio recordings, the larger model made the fewest transcription errors overall, and it matched the best speaker-identification results in 12 out of 13 test situations. Accuracy isn't perfect, but this is a strong step toward tools that understand conversation as quickly as humans hear it.

The best part for everyday use? The researchers aren't keeping it secret. They've released both the model weights—the actual trained brain of the AI—and the code to run it, free online. That means app developers, startups, and even hobbyists could start building better live transcription tools sooner than later. Within a few years, your meeting software might not just write down what was said; it'll tell you who said it, live, without breaking a sweat.

Key Points
  • The AI transcribes speech and identifies speakers at the same time, in real time, rather than in two slow steps.
  • The larger 7B model had the lowest transcription error across five test sets and near-best speaker labeling on 12 of 13 tests.
  • Both model versions and code are released publicly, so developers can start building live transcript tools that know who's talking.

Why It Matters

Live meeting notes, phone call captions, and voice assistants will get faster and more accurate—helpful at work, in class, and at home.

📬 Get the top 10 AI stories daily