New AI Builds a Huge Free Voice Dataset for Persian Speakers
Could mean Siri and Alexa finally speak Farsi properly.
Building voice AI for low-resource languages is hard: large-scale, high-fidelity speech datasets are scarce, traditional alignment-based methods need rare verbatim transcripts, and standard in-the-wild pipelines often rely on single-model automatic speech recognition and silence-based segmentation — leading to transcription errors and truncated prosody. PersianVox addresses this for Persian with a fully automated pipeline that generates high-quality speech corpora from unlabeled web data, yielding a 2,400-hour multi-speaker dataset described as the largest open-source speech resource available for Persian to date. It combines a prosody-aware segmentation strategy that uses acoustic turn-detection to preserve linguistic completeness and optimize utterance duration for long-context modeling, with a dual-model agreement mechanism that leverages two distinct model architectures to filter unreliable transcriptions without ground truth. The paper also provides the first comparative benchmark of speech quality assessment methods for Persian, releasing a human-annotated subset to facilitate future research.
- PersianVox turned raw web audio into 2,400 hours of Persian speech — the largest free dataset of its kind.
- Two AI models check each other's transcriptions, so no human has to type out a single word.
- The same recipe could bring decent voice AI to hundreds of smaller languages that tech giants ignore.
Why It Matters
Could bring working voice assistants, dictation and audiobooks to 110 million Persian speakers — and eventually many other languages.