Research & Papers

Top LLMs beat human examiners in GCSE grading accuracy study

⚑Study of 32,534 double-marked GCSE exams finds AI outmatches human markers.

Deep Dive

Researchers introduced a dataset of 32,534 double-marked real GCSE mock exam responses across 328 questions in five subjects, including handwritten work. They evaluated whether leading large language models agreed with examiners as closely as the two human examiners agreed with each other. Remarkably, the top-performing models agreed more closely with the examiner consensus than the human markers agreed with each otherβ€”a significant milestone for AI in education.

The models demonstrated strong performance on subjective tasks like English essay grading and handled complex, messy handwritten Maths papers effectively. Agreement was uniform near the pass/fail boundary, and performance was not heavily influenced by model size, indicating that smaller, cost-efficient models could be deployed for automated marking. This suggests a viable path toward scalable, accurate, and affordable AI-assisted grading for high-stakes national exams.

Key Points
  • Dataset: 32,534 double-marked real GCSE responses across 328 questions in 5 subjects, including handwritten work.
  • Top LLMs agreed more closely with human examiner consensus than the two human examiners agreed with each other.
  • Models handled subjective English essays and messy handwritten Maths; performance did not heavily depend on model size.

Why It Matters

AI grading could reduce cost and bias in high-stakes exams, matching or exceeding human accuracy.

πŸ“¬ Get the top 10 AI stories daily