Research & Papers

New study: LLM brand votes are 98% noise — language, not brand, drives answers

A variance decomposition of 12,933 LLM responses shows language explains 26.5% of brand-score variance.

Deep Dive

A new paper by Dmitrij Żatuchin takes a hard statistical look at why LLMs give inconsistent brand recommendations. The study, covering 12,933 responses across 20 Central and Eastern European brands, 8 languages, and 3 models (GPT-5.2, Gemini 3 Flash, Perplexity), uses a crossed random-effects decomposition to partition noise into four sources: within-prompt resampling, prompt paraphrase, model identity, and query language. The results are sobering for anyone using LLMs as brand sentiment tools.

Query language is the largest systematic facet, accounting for 26.5% of the variance in a single response, while brand identity contributes just 1.5% (ICC 0.0146). Resampling noise (34.8%) and brand-in-context interaction (29.6%) dominate. Per unit of query budget, adding languages and models reduces relative-error variance far more than adding repeats—a fifth repeat reduces error by only 0.0003. Brand-ranking reliability remains near 0.01 for a single answer and only 0.36 in the full crossed design, meaning reliable measurement requires spreading queries across languages and models, not repeating the same prompt.

Key Points
  • Language accounts for 26.5% of variance, brand identity only 1.5% (ICC 0.0146) — a single LLM answer carries almost no brand signal.
  • Resampling noise is 34.8% and brand-in-context interaction 29.6%; adding a fifth repeat reduces error by only 0.0003.
  • Reliability reaches just 0.36 in a full crossed design; spreading budget across languages and models is far more effective than repeating prompts.

Why It Matters

Marketers and researchers relying on LLM brand scores must redesign measurement to prioritize language and model diversity over repeated queries.

📬 Get the top 10 AI stories daily