AI Safety

EduClaw-Bench: No AI tutor sustains quality over 30 days

55 simulated tutoring scenarios reveal long-horizon agents still fail after days of interaction

Deep Dive

EduClaw-Bench, from researchers including Unggi Lee and colleagues, tackles a blind spot in AI tutoring: most LLM educational tools are point solutions for single tasks, like essay scoring or one-turn Q&A. But real tutoring is long-horizon—students improve over days or weeks. The benchmark places an agent tutor in a continuous 30-day relationship with a simulated learner. The learner's knowledge-concept mastery comes from a knowledge tracing (KT) model trained on real-student data, which drives its answers and enables measurement of learning gain across 55 scenarios. Agents are scored on three primary axes (learning gain, responsiveness, helpfulness) and two curriculum-design axes (Gagné and Rosenshine), with the latter judged by a panel of three LLM judges from different model families.

Evaluating 10 agent adapters over three base-model tiers, the authors reach two findings that single-session benchmarks miss. First, tutoring quality is a joint property of the base model and the agent harness—not just one or the other. Second, almost no combination maintains good tutoring over the full 30-day horizon, suggesting current systems degrade as the simulated relationship progresses. A calibration check (ECE=0.049) and a live-classroom field study confirm the simulated learner and its measurements track reality. This is a step toward trustworthy AI tutors for future education, but it also warns that we're far from deployable long-term AI teaching agents.

Key Points
  • EduClaw-Bench simulates a 30-day tutor-student relationship using knowledge tracing trained on real-student data across 55 scenarios
  • 10 agent adapters on 3 base-model tiers were evaluated; almost no combination sustained quality over the full horizon
  • Calibration ECE=0.049 and a live-classroom field study confirm the simulated learner tracks real-world outcomes

Why It Matters

Long-horizon AI tutoring needs both strong base models and careful agent design—single-turn benchmarks overstate real-world teaching quality.

📬 Get the top 10 AI stories daily