AI Safety

AI Can Now Write the Test and Grade It, Study Finds

Teachers spend hours grading written work. This AI does it instantly — and students can't tell.

Deep Dive

A team of Japanese researchers has shown that a large language model — the same kind of AI behind ChatGPT — can take over one of the most time-consuming jobs in professional education: creating practice problems with deliberate mistakes in them, grading students' written answers, and explaining what they got wrong. The teaching method, called hierarchical diagnostic reasoning, asks learners to read a realistic case (a hospital mix-up, a flawed engineering plan) and identify and explain the errors. It's excellent at building judgment, but it's expensive, because experts must hand-write every case and read every answer.

The surprising part is how little setup it took. The researchers didn't retrain or customize the AI with piles of specialist data — they simply wrote good instructions, a technique known as prompt design. That means a nursing school or a management training program could adopt this without hiring data scientists. When the AI's grades were compared with human instructors' grades, they matched completely under some conditions, and the AI's written feedback was rated by students as just as convincing and useful as feedback from a real teacher.

The researchers also checked whether the AI-generated problems were fair and consistent, using a standard statistical measure called Cronbach's alpha, which scores how reliably a set of questions measures the same skill. It came in at 0.78 — a respectable result suggesting the questions hang together well rather than being random.

So what does this mean for you? If you've ever waited weeks for feedback on a certification exam, a work assessment, or a professional course, this points to a future where you get detailed comments within seconds, and where you can be re-tested repeatedly over months to track real improvement instead of a one-off score. The catch: this is a single research study, not a product you can buy, and 'matched human graders under some conditions' is not the same as 'always matches.' AI graders can still be confidently wrong, and any high-stakes decision about a person's career should keep a human in the loop.

Key Points
  • The AI handled all three jobs — writing the practice test, grading written answers, and explaining mistakes — with no special training, just well-written instructions.
  • In some tests, its grades matched human teachers exactly, and students rated its feedback as useful as a real instructor's.
  • The likely payoff is speed and cost: professional training that once took experts weeks to build and grade could run repeatedly and almost instantly.

Why It Matters

Faster, cheaper professional training and near-instant feedback could reach people who never had access to expert tutors.

📬 Get the top 10 AI stories daily