Research & Papers

PatiGonit22K dataset boosts Bengali math reasoning with 22K problems

Bengali math word problems get a massive upgrade with 22,441 culturally-adapted questions.

Deep Dive

Mathematical Word Problems (MWPs) are a key benchmark for natural language understanding and quantitative reasoning, yet Bengali has remained underserved due to a lack of large-scale annotated datasets. To address this, a team of researchers—Swastika Kundu, Azizul Hakim Fayaz, and Tashreef Muhammad—released PatiGonit22K, a comprehensive dataset of 22,441 Bengali MWPs. Built by extending the original PatiGonit dataset, it includes a substantially larger collection of complex problems spanning both simple and multi-operation equations, providing a balanced benchmark across difficulty levels. Each problem was meticulously translated, annotated, culturally adapted, and verified to ensure linguistic consistency and mathematical correctness.

By increasing both the scale and complexity of Bengali MWPs, PatiGonit22K offers a more comprehensive resource for advancing research in mathematical reasoning and educational NLP applications for low-resource languages. This dataset enables researchers and educators to train and evaluate AI models on culturally relevant math problems, potentially impacting millions of Bengali speakers. The work, available on arXiv (2607.22859), represents a significant step toward closing the language gap in mathematical AI benchmarks and fostering inclusive AI development in education.

Key Points
  • Dataset contains 22,441 Bengali mathematical word problems, extending the original PatiGonit dataset.
  • Includes both simple and multi-operation equations to benchmark reasoning across difficulty levels.
  • Each problem is carefully translated, culturally adapted, and verified for linguistic and mathematical accuracy.

Why It Matters

Bridges a critical gap for Bengali NLP, enabling AI-driven math education tools for over 300 million speakers.

📬 Get the top 10 AI stories daily