AI Safety

mmPISA-bench: LLMs Reason Equally Well Across 43 Languages

New benchmark shows LLMs match human accuracy in 43 languages with machine translations.

Deep Dive

A new research paper by Yerzhan Sapenov and Jaromir Savelka presents mmPISA-bench, a compact, high-quality multilingual reasoning benchmark derived from the OECD Programme for International Student Assessment (PISA). The benchmark consists of 25 carefully selected multiple-choice questions that require genuine reasoning to answer correctly. Each question is available in official human translations for 43 languages, supplemented by machine-translated versions, yielding a total of 2,150 data points. The researchers evaluated two mainstream proprietary large language models (LLMs) across all languages, varying reasoning effort levels and translation types to assess their accuracy.

The results are striking: modern LLMs can reason effectively across all 43 languages tested, achieving accuracy comparable to human test-takers. Importantly, performance variations between languages are minimal, indicating strong multilingual generalization. The study also found that using high-quality machine translations does not degrade accuracy relative to official human translations, suggesting that synthetic data is often sufficient for large-scale multilingual reasoning evaluations. However, the analysis of token usage and inference cost reveals a notable disparity: some languages are simultaneously more expensive and less accurate, highlighting economic inequities in current LLM deployment. Overall, mmPISA-bench provides a robust, scalable tool for assessing multilingual reasoning and underscores the viability of machine translation for evaluation pipelines.

Key Points
  • LLMs achieve accuracy comparable to human test-takers across all 43 languages on the PISA-derived benchmark.
  • Machine-translated questions show no significant accuracy loss compared to official human translations.
  • Inference cost varies by language, with some languages being both more expensive and less accurate.

Why It Matters

Enables reliable multilingual AI evaluation without expensive human translations, highlighting cost-performance trade-offs across languages.

📬 Get the top 10 AI stories daily