Research & Papers

New study: GPT-5 and Claude 4 match human graders in math and science

Claude Sonnet 4 and GPT-5 achieve near-human grading accuracy on MCAS tests.

Deep Dive

A study presented at the NCME 2026 Conference evaluated four commercially available foundational models—Claude Sonnet 4, Haiku 4.5, GPT-5, and GPT-5 Mini—as K-12 assessment graders through context and prompt engineering. Using Massachusetts Comprehensive Assessment System (MCAS) data, researchers measured interrater agreement via Quadratic Weighted Kappa (QWK) and Proportional Reduction in Mean-Squared Error (PRMSE) across mathematics, science, and English language arts (ELA). The results demonstrated that LLM graders, particularly those with more parameters like Claude Sonnet 4 and GPT-5, achieved substantial agreement with human raters in mathematics and science, validating the potential for AI-driven standards-based grading in these subjects.

However, performance varied significantly in ELA assessments, where the subjective nature of writing evaluation posed challenges for generic foundation models. Additional analysis of teacher and student feedback revealed strong acceptance of AI-generated narrative feedback but skepticism toward numerical scores, indicating that educators trust AI for formative insights but not for summative evaluation. The authors advocate for hybrid models that combine AI efficiency with teacher judgment, claiming such systems can reduce workload, enhance feedback quality, and support equitable assessment practices without displacing professional expertise.

Key Points
  • Tested Claude Sonnet 4, Haiku 4.5, GPT-5, and GPT-5 Mini for grading K-12 assessments using MCAS data.
  • Achieved high interrater agreement (QWK/PRMSE) with human raters in math and science; results varied in ELA.
  • Teachers and students favored AI narrative feedback but distrusted AI-generated numerical scores.

Why It Matters

Hybrid AI-teacher grading could reduce educator workload and improve feedback quality in K-12 education.

📬 Get the top 10 AI stories daily