Research & Papers

MLLMs fail as synthetic study participants: Gemini 3, Qwen 3 show bias

AI can't replicate human video ratings—new study reveals hidden biases.

Deep Dive

Researchers evaluated multimodal LLMs (Gemini 3 Flash, Qwen 3 Omni) as synthetic participants in video-based sensory engagement studies. Using the Perceived Message Sensation Value framework, they compared ratings from human participants and profile-conditioned MLLM simulations (n=673) with a 17-item scale measuring emotional arousal, dramatic impact, and novelty. Results showed limited agreement, with downward mean-shift and central-tendency biases. Models introduced and flattened subgroup differences while responding inconsistently to participant profiles. Prompting strategies affected metrics differently, modestly improving some aspects while worsening others. The study highlights both challenges and opportunities of using MLLMs as synthetic participants in video-based research.

Key Points
  • Gemini 3 Flash and Qwen 3 Omni both showed poor agreement with human ratings on a 17-item sensory engagement scale.
  • Models exhibited systematic downward mean-shift and central-tendency biases, compressing rating distributions.
  • Prompting strategies improved some metrics but worsened others, with no reliable way to simulate participant profiles.

Why It Matters

MLLMs can't yet replace human subjects—AI biases distort subjective research, requiring caution in HCI studies.

📬 Get the top 10 AI stories daily