Research & Papers

80% of Mechanical Turk surveys use AI, study finds

LLM-assisted survey responses range from 10% to over 80% by platform

Deep Dive

A new study by Zane Xu and Nathan Malkin, presented at SOUPS 2026, quantifies how often survey respondents on crowdsourcing platforms use large language models (LLMs) to generate answers. Across 250 surveys, they tested conditions such as platform (Prolific vs. Amazon Mechanical Turk), survey length, requests to avoid AI, and disabling copy-paste. Results showed a stark contrast: fewer than 10% of responses on Prolific showed signs of LLM assistance, while over 80% on Mechanical Turk did.

Mitigation efforts (e.g., disabling copy-paste) reduced LLM usage but paradoxically didn't improve data quality, suggesting respondents may still cheat via other means. No participants used browser-based AI agents, but the authors conducted their own experiments to detect such automation. The paper recommends active screening: recording keystroke dynamics and designing instructions and questions that specifically deter or identify AI-generated content. The findings highlight a growing threat to the validity of behavioral research relying on crowdsourced data.

Key Points
  • LLM-assisted survey responses ranged from under 10% (Prolific) to over 80% (Mechanical Turk).
  • Disabling copy-paste reduced AI usage but did not improve overall data quality.
  • Researchers recommend keystroke recording and AI-specific prompts to detect AI answers.

Why It Matters

Crowdsourced research data may be silently contaminated by AI, threatening the validity of thousands of studies.

📬 Get the top 10 AI stories daily