60M-line religious radio transcript corpus opens new research frontiers
Over 700,000 recordings from 2,000+ stations, all transcribed and labeled by AI.
A corpus of transcribed English-language religious radio broadcasts from 785 webstreams over July 2025, capturing over 2,000 AM/FM stations. The dataset includes over 700,000 recordings and over 60 million speaker-diarized lines, transcribed via automated pipelines and labeled by format/topic using a large language model. It enables analysis of regional religious broadcasting, socio-political discourse in religious media, and speech-processing research.
- 785 live webstreams captured over July 2025, rebroadcasting 2,000+ AM/FM stations
- 700,000+ 15-minute recordings yielding 60+ million diarized lines of speech
- Automated transcription and LLM-based labeling by format and topic for analysis
Why It Matters
Unlocks large-scale content analysis of religious radio, a major but understudied media sector.