BENI Global 10 Corpus Bridges Economic Narrative Gap for Global South with 10 Languages
84% of economic sentiment research ignores the Global South – until now.
Economic narrative indices have long been dominated by English-language data, with 84% of sentiment-based forecasting research focusing on developed economies. BENI Global 10, created by Ann Naser Nabil, directly tackles this bias by providing the first multilingual economic news corpus designed specifically for the Global South. The dataset spans 10 languages across 7 language families and 5 economic regions: Bangladesh (Bangla), India (Hindi), Turkey (Turkish), Indonesia (Indonesian), Brazil (Portuguese), Egypt (Arabic), Vietnam (Vietnamese), Philippines (Filipino), Kenya (Swahili), and Pakistan (Urdu).
To build the corpus, over 2.8 million raw documents were collected and filtered using 25–32 translated economic keywords per language, resulting in 522,397 relevant articles. The release includes a reproducible streaming pipeline with checkpoint-resume for low-resource environments, per-language schema-normalized Parquet files with economic relevance labels, and a temporally synced cross-lingual index covering 2018–2024. Inter-annotator agreement was validated with Cohen's kappa > 0.70 across all languages. The complete dataset, code, and annotation guidelines are publicly available for research use, opening up new opportunities for economic forecasting and policy analysis in underrepresented regions.
- 522,397 economically relevant articles across 10 languages from 5 Global South regions
- Filtered from 2.8M raw documents using 25–32 translated keywords per language
- Inter-annotator agreement exceeds kappa > 0.70; full dataset, code, and guidelines released publicly
Why It Matters
Enables AI-driven economic forecasting and sentiment analysis for billions of people in previously understudied languages and regions.