Research & Papers

LLMs can generate synthetic consumer data matching human responses

New study reveals AI-generated consumer insights rival real survey data on broad topics.

Deep Dive

A new study on arXiv tests whether large language models (LLMs) can replace costly, time-intensive consumer surveys by generating synthetic insight data. Authors France and Albinsson used five projective marketing tasks—word association, sentence completion, storytelling, etc.—with multiple LLMs (including variants of GPT), varied prompting strategies, and different temperature settings. They compared LLM outputs against human responses from a primary study on perceptions of city tourism destinations.

Analyzing the data with linguistic measures, diversity/concentration metrics, topic models, and top-term analysis, the researchers found high overlap in broad themes and associations between human and LLM responses. However, significant differences emerged in writing style, linguistic structure, and how diversity is generated. LLMs tended to produce more uniform, less varied language, while humans offered richer idiosyncratic detail. The study provides actionable guidance on model selection, prompt design, and temperature tuning to maximize the usefulness of synthetic data, while cautioning against overreliance due to these stylistic biases.

Key Points
  • Tested multiple LLMs (including GPT variants), prompting strategies, and temperature settings across five projective marketing tasks.
  • Found substantial overlap in broad topics and associations between human and LLM responses on city tourism perceptions.
  • Identified key differences: LLMs produced less diverse language and different stylistic/linguistic structures than humans.

Why It Matters

LLMs could slash costs and accelerate consumer research, but marketers must account for stylistic and diversity biases.

📬 Get the top 10 AI stories daily