MLLMs fail as synthetic study participants: Gemini 3, Qwen 3 show bias
AI can't replicate human video ratings—new study reveals hidden biases.
Researchers evaluated multimodal LLMs (Gemini 3 Flash, Qwen 3 Omni) as synthetic participants in video-based sensory engagement studies. Using the Perceived Message Sensation Value framework, they compared ratings from human participants and profile-conditioned MLLM simulations (n=673) with a 17-item scale measuring emotional arousal, dramatic impact, and novelty. Results showed limited agreement, with downward mean-shift and central-tendency biases. Models introduced and flattened subgroup differences while responding inconsistently to participant profiles. Prompting strategies affected metrics differently, modestly improving some aspects while worsening others. The study highlights both challenges and opportunities of using MLLMs as synthetic participants in video-based research.
- Gemini 3 Flash and Qwen 3 Omni both showed poor agreement with human ratings on a 17-item sensory engagement scale.
- Models exhibited systematic downward mean-shift and central-tendency biases, compressing rating distributions.
- Prompting strategies improved some metrics but worsened others, with no reliable way to simulate participant profiles.
Why It Matters
MLLMs can't yet replace human subjects—AI biases distort subjective research, requiring caution in HCI studies.