Top LLMs beat human examiners in GCSE grading accuracy study
Study of 32,534 double-marked GCSE exams finds AI outmatches human markers.
Researchers introduced a dataset of 32,534 double-marked real GCSE mock exam responses across 328 questions in five subjects, including handwritten work. They evaluated whether leading large language models agreed with examiners as closely as the two human examiners agreed with each other. Remarkably, the top-performing models agreed more closely with the examiner consensus than the human markers agreed with each otherβa significant milestone for AI in education.
The models demonstrated strong performance on subjective tasks like English essay grading and handled complex, messy handwritten Maths papers effectively. Agreement was uniform near the pass/fail boundary, and performance was not heavily influenced by model size, indicating that smaller, cost-efficient models could be deployed for automated marking. This suggests a viable path toward scalable, accurate, and affordable AI-assisted grading for high-stakes national exams.
- Dataset: 32,534 double-marked real GCSE responses across 328 questions in 5 subjects, including handwritten work.
- Top LLMs agreed more closely with human examiner consensus than the two human examiners agreed with each other.
- Models handled subjective English essays and messy handwritten Maths; performance did not heavily depend on model size.
Why It Matters
AI grading could reduce cost and bias in high-stakes exams, matching or exceeding human accuracy.