European MMLU Dataset Brings LLM Benchmark to 11 Languages
A new multilingual MMLU dataset aims to level the AI evaluation playing field.
Researchers from the DGT and EMT network have completed a project to localize the widely-used MMLU (Massive Multitask Language Understanding) dataset into 11 European languages. MMLU tests LLMs across 57 subjects, from law to physics, but until now was available mostly in English. This localization — covering languages like French, German, Spanish, Italian, Portuguese, Dutch, Polish, Swedish, Czech, Greek, and Bulgarian — enables fairer evaluation of AI models for non-English users. The project involved master's students in translation and project management, who gained real-world experience in multilingual coordination, revision, and workflow management, alongside technical challenges like handling domain-specific terminology and ensuring consistency across language pairs.
The initiative highlights significant methodological hurdles, such as maintaining semantic equivalence across diverse linguistic structures and dealing with cultural references that don't translate directly. Administrative challenges included aligning academic calendars with project timelines and coordinating across institutions in different countries. The resulting dataset is expected to become a key resource for evaluating LLMs in European contexts, supporting more accurate model performance in languages outside English. It also demonstrates a scalable model for future localization of AI benchmarks, combining educational goals with practical AI infrastructure building. The paper was presented at the 3rd International Conference on New Trends in Translation and Interpreting Technology in Dubrovnik.
- Localized MMLU into 11 European languages: French, German, Spanish, Italian, Portuguese, Dutch, Polish, Swedish, Czech, Greek, and Bulgarian.
- Partnership between the European Commission's DGT and the EMT network, involving real-world training for master's students.
- Covers 57 subjects from law to physics, aiming to create a more inclusive benchmark for LLM evaluation across languages.
Why It Matters
This dataset enables fairer, more diverse LLM evaluation in European languages while training future translators on real AI tasks.