AI Safety

CDP technique makes LLM-simulated examinees 90%+ aligned with humans

Gemini 3.0 Flash boosts item-difficulty correlation from 0.24 to 0.90

Deep Dive

Psychometric calibration for educational tests traditionally requires expensive human response data. Large language models (LLMs) can simulate examinees cheaply, but their responses are too accurate and uniform to represent real students. To bridge this gap, researchers from the University of California, Berkeley and affiliated institutions developed Cognitive Diagnostic Profiling (CDP), a zero-shot framework that prompts LLMs to generate plausible examinees with varied cognitive profiles. CDP works by converting binary attribute-mastery patterns into natural-language descriptions, then sampling these patterns from either an uninformative or informative distribution. This approach allows LLMs to exhibit realistic diversity in ability levels and item response patterns without any fine-tuning.

Evaluating eight LLM configurations on the classic Tatsuoka fraction-subtraction dataset, CDP significantly improved alignment with human examinees at three levels: ability distribution, mastery-profile scores, and item difficulty. The informative CDP variant yielded weighted profile-level correlations of 0.92 to 0.98. Most notably, with Google's Gemini 3.0 Flash (Thinking) model, the one-parameter logistic difficulty Spearman correlation jumped from 0.24 (no profile) to 0.86 (uninformative CDP) and 0.90 (informative CDP). Simultaneously, root-mean-square error (RMSE) fell from 6.31 to 1.30 and finally 0.90. These results demonstrate that CDP can make LLM-simulated examinees practically indistinguishable from human data for calibration purposes, potentially saving education departments and testing organizations millions of dollars in data collection costs while accelerating test development cycles.

Key Points
  • CDP improved profile-level correlations to 0.92–0.98 across tested LLM configurations
  • Gemini 3.0 Flash (Thinking) item-difficulty Spearman rose from 0.24 to 0.90 under informative CDP
  • Root-mean-square error dropped from 6.31 to 0.90, making simulated data viable for operational calibration

Why It Matters

Enables low-cost, scalable psychometric calibration for educational tests using realistic LLM-simulated examinees.

📬 Get the top 10 AI stories daily