Khan Academy's Khanmigo paper reveals how to improve AI tutoring quality
Khanmigo's 15-page AIED paper shares experimental methods for boosting tutor engagement.
Khan Academy, the nonprofit education platform, released a research paper titled "Methodologies for Improving the Quality of AI Tutoring in K-12 Education," accepted at the 27th International Conference on Artificial Intelligence in Education (AIED 2026). Authored by Tushar Udeshi and six colleagues, the paper draws on real-world experience from Khanmigo, Khan Academy's AI-powered tutor launched in 2023. The researchers argue that because LLMs are opaque, robust evaluation and live experimentation are critical for every change. They detail the specific metrics they use to assess tutoring quality and student engagement, and they share the experiments that led to measurable improvements.
Key improvements came from adjustments across four areas: model selection, prompt engineering, personalization, and AI agents. The paper highlights which changes moved the metrics, offering concrete evidence for what works in AI tutoring. For example, the team tested different LLMs and prompting strategies to reduce hallucinations and improve pedagogical responses. By iterating on personalization and agent-based features, they boosted engagement and learning outcomes. The paper is published in the Springer LNCS volume (vol. 16582) and is available on arXiv under DOI 10.48550/arXiv.2608.11259. This work provides a rare behind-the-scenes look at how a major education platform systematically improves an AI tutor, making it a valuable reference for both researchers and edtech practitioners.
- Khan Academy's 15-page paper accepted at AIED 2026 details quality metrics for AI tutoring and engagement measurement.
- Experiments focused on four levers: model selection, prompting, personalization, and agents—each showing measurable impact.
- Published as arXiv:2608.11259 and in Springer LNCS vol. 16582, giving researchers an open methodology for AI tutor refinement.
Why It Matters
This gives edtech teams a proven framework for measuring and improving AI tutoring quality in real classrooms.