Research & Papers

New arXiv study finds Qwen3 and GPT-5-Mini can't mimic human belief shifts

Six LLMs failed to simulate belief updates from 391 human participants, study finds.

Deep Dive

A new arXiv paper (2607.28347) from Sebastian Pohl and colleagues tests whether LLMs can act as stand-ins for human participants in social science research. The team compared six models — including Qwen3-32B and GPT-5-Mini — against belief updates from 391 UK participants on Prolific, who read Reddit comments and shifted their stances on three topics. Each LLM was given a persona based on demographic and personality data, then asked to simulate the individual's initial stance and subsequent belief change. Only Qwen3-32B and GPT-5-Mini matched the human post-stance distribution, and only when they were fed participants' actual starting positions. No model successfully simulated initial stances from persona data alone, and all failed to produce faithful belief updates from self-generated stances.

Three systematic biases appeared across every model: overrepresentation of neutral positions, belief shifts that were more frequent but smaller than real humans, and an inability to rank comments by convincingness. Adding demographic and personality personas had no consistent effect on simulation fidelity. The authors conclude that LLM-based simulations of belief dynamics are only reliable when grounded in realistic starting conditions — conditions that current multi-round social media simulations rarely provide. The study directly challenges the growing practice of using LLMs as proxy participants in social science experiments and offers a concrete benchmark (with 391 human-validated data points) for evaluating future models.

Key Points
  • Only Qwen3-32B and GPT-5-Mini matched human post-stance distributions, and only with actual initial stances provided
  • All six models failed to simulate initial stances from demographic and personality personas alone
  • Three consistent biases: overrepresentation of neutral positions, smaller/more numerous belief shifts, and poor ranking of comment convincingness

Why It Matters

LLM-as-participant research is unreliable; social science simulations need real starting stances to produce valid results.

📬 Get the top 10 AI stories daily