Research & Papers

New Study Benchmarks LLMs' Ability to Generate Creative Educational Questions

20,700 questions across subjects reveal which LLMs truly stimulate higher-order thinking.

Deep Dive

The paper "From Memorization to Creation" evaluates six widely-used LLMs through the lens of Bloom's Taxonomy, focusing on their capacity to generate questions that go beyond simple recall. The researchers, from multiple institutions and accepted at KDD 2026, created a dataset of 20,700 questions spanning computer science, K–12 math, and social sciences. Using a hybrid human-AI evaluation protocol, they assessed cognitive depth. They introduced a fine-grained prompting strategy that reduced question repetitiveness by 24.45% for Qwen2.5-7B-Instruct and increased the proportion of higher-order cognitive level outputs by 11.53% for InternLM3-8B-Instruct. The study also proposes quantitative metrics: CogShift (cognitive shift intensity) and category drift to measure multi-level transitions.

The analysis revealed that InternLM3-8B-Instruct outperformed other models in achieving multi-level cognitive transitions, suggesting that certain architectures are more amenable to generating higher-order thinking questions. An interpretability analysis highlighted metric-level correlations that enhance transparency when using Chain-of-Thought prompting. The findings underscore the importance of cognitive-aware prompt design for deploying LLMs in personalized learning systems. The paper provides benchmarks for future work and demonstrates that with the right prompting, LLMs can transcend rote memorization to create questions that stimulate analysis, evaluation, and creation. This has implications for automated tutoring systems, assessment generation, and adaptive learning platforms.

Key Points
  • Generated 20,700 questions across CS, K-12 math, and social science using six different LLMs.
  • Prompt engineering reduced question repetitiveness by 24.45% (Qwen2.5-7B-Instruct) and increased higher-order cognitive outputs by 11.53% (InternLM3-8B-Instruct).
  • InternLM3-8B-Instruct outperformed peers in multi-level cognitive transitions, measured by the new CogShift metric.

Why It Matters

Personalized learning systems can now leverage cognitive-aware prompts to generate deeper, more engaging questions.

📬 Get the top 10 AI stories daily