LLM survey respondents fail psychometric audit: plausible but not valid
37 LLMs from OpenAI, Anthropic, Google tested — a simple statistical model beats them all.
Researchers Mantas Lukauskas and Viktorija Šarkauskaitė conducted a comprehensive psychometric audit of LLMs as synthetic survey respondents, testing 37 models from OpenAI, Anthropic, Google, and twelve open-weight families. Using a Lithuanian organizational-psychology dataset (n=263 employees, 68 items across 12 subscales), they applied a five-level persona-disclosure ladder, demographic counterfactual swaps, and cross-language checks. The headline finding: LLMs reproduce the qualitative direction of human responses but fail psychometric validity — a simple Gaussian-copula statistical baseline beat every LLM on sample-driven similarity metrics.
The audit introduces a Psychometric Similarity Score (PSS) anchored against human-vs-human ceilings. LLMs scored 0.73 inter-LLM similarity (crowds are more like themselves than humans), memorization didn't drive performance (recall-PSS correlation 0.00), and demographic swaps revealed education-driven effects (mean |d|=0.56) far exceeding gender (0.12) and role (0.18). Downstream, all LLMs exhibited a strong acquiescence shift (+0.84 SD), synthetic-trained regressors lost validity on held-out humans (mean R² dropped from 0.28 to -0.18), and models fabricated indirect effects on 3 of 10 placebo mediation paths. The authors conclude LLM samples are not a drop-in replacement for human survey data.
- Gaussian-copula baseline beats all 37 LLMs on psychometric similarity
- Inter-LLM PSS 0.73 — LLM crowds more similar to themselves than to humans
- Synthetic-trained models lose predictive validity: R² drops from 0.28 to -0.18
Why It Matters
Researchers and marketers using LLM-generated survey responses risk statistically invalid insights — a simple classical model may be more trustworthy.