AI Safety

ChatGPT 5.5 Thinking tops LLM math grading study with liberal prompts

Open-ended math exams graded with 1.87 MAE error – rivaling human TAs?

Deep Dive

Grading open-ended math exams at scale is notoriously difficult: instructors must apply partial-credit rubrics consistently while giving actionable feedback. A new arXiv paper (2607.01247) by Sapkota and Murshed tested six contemporary LLMs as automated teaching assistants for an undergraduate discrete mathematics exam. The models included Gemini 3.1 Pro Extended, Gemini 3.5 Flash, ChatGPT 5.5 Pro Extended, ChatGPT 5.5 Thinking, Claude Pro Opus 4.7, and Claude Sonnet 4.6. Each was run under two grading policies: a strict BASELINE prompt emphasizing explicit evidence, and a LIBERAL prompt that allowed more lenient partial credit.

The results show that liberal partial-credit prompting reduces average question-level error for every model family. ChatGPT 5.5 Thinking under the LIBERAL condition had the lowest question-level MAE (1.87) and RMSE (2.53), meaning its per-question grading deviations were the smallest. For total exam scores, Gemini 3.1 Pro Extended (LIBERAL) yielded the lowest MAE (8.00) and RMSE (10.66). However, the highest Pearson correlation between LLM and human total scores was only 0.58, achieved by Gemini 3.1 Pro Extended under the BASELINE prompt—indicating that while some models calibrate absolute scores well, rank-ordering of students remains a distinct challenge.

Practical usability observations from the study highlight that LLMs can handle rubric-based grading but struggle with implicit reasoning or creative proofs. The liberal policy occasionally over-credited vague answers but generally aligned better with human graders. The authors emphasize that LLMs are not yet a replacement for human examiners but can serve as effective first-pass checkers or consistency monitors. For institutions grading large STEM classes, this suggests a hybrid workflow: liberal LLM grading for speed, followed by human spot-checking of outlier scores.

Key Points
  • ChatGPT 5.5 Thinking with liberal prompting achieved the lowest question-level MAE of 1.87 and RMSE of 2.53.
  • Gemini 3.1 Pro Extended (liberal) had the best total-score MAE of 8.00, but its baseline prompt gave the highest Pearson correlation (0.58).
  • Liberal rubric prompts reduced error for all six tested models compared to strict rubric-following prompts.

Why It Matters

Reliable LLM grading could save thousands of TA hours in large STEM courses while maintaining consistency.

📬 Get the top 10 AI stories daily