GeoDial dataset brings visual tutor turns to geometry tutoring AI
1,300+ teacher-student dialogs with diagram highlights to train better AI tutors.
A team of researchers from ETH Zürich and other institutions has released GeoDial, a novel multimodal conversational tutoring dataset designed to advance AI tutors for geometry problem-solving. Unlike most existing tutoring datasets that rely solely on text, GeoDial explicitly grounds instructional turns in diagram highlights, mirroring how human teachers use visual cues. The dataset comprises over 1,300 teacher-student dialogs collected from experienced math teachers, accompanied by a scalable annotation protocol that integrates dialog acts, visual highlighting, and feedback.
To benchmark current capabilities, the authors fine-tuned several vision-language models on GeoDial. While supervised fine-tuning substantially improved the quality of generated tutoring utterances, the models struggled to produce accurate diagram highlights—a critical component for effective visual instruction. This limitation underscores the need for new approaches that more effectively combine visual reasoning with pedagogical interaction. GeoDial provides a valuable testbed for developing AI systems that can teach geometry with the same visual-grounded techniques used by human instructors.
- Dataset includes over 1,300 teacher-student geometry dialogs with diagram highlights for each tutor turn.
- Fine-tuning vision-language models on GeoDial improved dialogue quality but failed to generate accurate diagram highlights.
- GeoDial uses a scalable annotation protocol combining dialog acts, visual highlighting, and feedback for fine-grained supervision.
Why It Matters
Enables development of AI tutors that teach geometry with visual grounding, a critical gap in current educational AI.