Research & Papers

60M-line religious radio transcript corpus opens new research frontiers

Over 700,000 recordings from 2,000+ stations, all transcribed and labeled by AI.

Deep Dive

A corpus of transcribed English-language religious radio broadcasts from 785 webstreams over July 2025, capturing over 2,000 AM/FM stations. The dataset includes over 700,000 recordings and over 60 million speaker-diarized lines, transcribed via automated pipelines and labeled by format/topic using a large language model. It enables analysis of regional religious broadcasting, socio-political discourse in religious media, and speech-processing research.

Key Points
  • 785 live webstreams captured over July 2025, rebroadcasting 2,000+ AM/FM stations
  • 700,000+ 15-minute recordings yielding 60+ million diarized lines of speech
  • Automated transcription and LLM-based labeling by format and topic for analysis

Why It Matters

Unlocks large-scale content analysis of religious radio, a major but understudied media sector.

📬 Get the top 10 AI stories daily