Indian AI models fall short on modern benchmarks, Sarvam AI leads
Indian models ace MMLU but skip agentic tests, new BMI framework finds.
Researchers Avinash Agarwal and Vridhi Jain have published a paper on arXiv (arXiv:2608.11891) titled "Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework". The study analyzes publicly benchmarked Indian AI models against global frontier and comparable-scale models across eight domains: reasoning, coding, agentic AI, cybersecurity, vision, video/multimodal understanding, scientific research, and Indic language capability. Using only public benchmark results, they found that Indian models perform strongly on established tests like MMLU and MATH-500, but these benchmarks are now considered saturated and frontier developers have largely stopped reporting them.
The authors introduce a four-dimension Benchmark Maturity Index (BMI) that scores each domain on standardization, participation, independent verification, and national coverage. Indian models participate far less in newer agentic and domain-specialized evaluations, and participation is uneven—Sarvam AI leads with the broadest coverage by a substantial margin. The BMI refines and sometimes revises maturity judgments derived from purely descriptive reviews. The key insight: many apparent capability gaps in the public record may actually be evaluation-ecosystem gaps. This distinction is crucial for how national AI programs design monitoring and funding criteria, as India and other governments increasingly fund indigenous models for digital sovereignty and multilingual computing.
- Indian models score well on saturated benchmarks (MMLU, MATH-500) but participate far less in newer agentic and domain-specialized evaluations.
- Sarvam AI reports the broadest benchmark coverage among Indian organizations by a substantial margin, per the paper's survey.
- The proposed Benchmark Maturity Index (BMI) uses four dimensions — standardization, participation, independent verification, and national coverage — to distinguish true capability gaps from evaluation-ecosystem gaps.
Why It Matters
Provides a structured framework for governments and investors to measure national AI progress — and avoid misreading evaluation gaps as capability failures.