AI Safety

AI Just Aced Teacher Certification Exams in Three Countries

⚡AI tutors are heading to classrooms — but they still flunk one key teaching skill.

Deep Dive

Researchers put 36 foundation-model variants through EDU 1.0, a benchmark of 10,012 questions taken from teacher certification and recruitment exams in the United States, China, and India — including the U.S. Praxis series, China's National Teacher Qualification Examination, and India's Kendriya Vidyalaya Sangathan examinations. The strongest proprietary model reached a response-balanced score of 96.5%. The leading open-weight model trailed by 2.1 percentage points, while the leading system deployable on a single accelerator reached 92.2%. Those aggregates conceal a shared limitation: all three systems scored higher on general pedagogical principles than on assessments requiring pedagogy to be applied within a discipline. Their subject-assessment scores spanned 4.1 to 8.8 points, and their shortfall relative to general pedagogy widened from 2.6 to 4.1 points as capability declined. The paper's conclusion is that the outstanding requirement is pedagogical content knowledge, the capacity to make particular subject matter teachable to particular learners, rather than general pedagogy or model scale.

Key Points
  • Researchers tested 36 AI models on 10,012 real teacher certification and hiring exam questions from the U.S., China, and India.
  • The best paid AI scored 96.5%, and a free model small enough to run on one chip still reached 92.2% — so strong teaching knowledge is now cheap.
  • Every model was weaker at applying teaching skill inside a specific subject, scoring 4 to 9 points lower than on general teaching theory.

Why It Matters

AI tutors are about to get cheap and everywhere — great for facts, still shaky at actually explaining tough ideas to your kid.

📬 Get the top 10 AI stories daily