Study reveals RLHF flattens diverse global AI preferences into one model
Only 49% want truthfulness, but even they can't agree on what it means.
A new study from arXiv (accepted at FAccT '26) systematically maps what people across the globe actually want from AI systems — and finds current alignment methods fundamentally inadequate. Authors Julia Sepúlveda Coelho and Scott A. Hale analyzed 1,500 open-ended responses from the PRISM dataset, covering 75 countries. Their core finding: preference plurality is real and deep. No single value is universally desired; truthfulness comes closest at 49%, but every other value falls below 25%. Worse, even when people use the same word — like 'truthfulness' — they mean different things: some ask for sourced claims, some for expert opinions, and some even demand that AI surface unpopular views. This makes a single reward model inherently incapable of capturing actual human preferences.
The paper also finds that certain capabilities are outright controversial. Traits like 'how human-like a model behaves' and specific features like AI guardrails divide users — some want them, others reject them. Moreover, people frequently use contextual distinctions that binary comparisons cannot capture, such as what AI should do 'by default' versus 'if requested.' The authors argue these findings expose fundamental problems in current RLHF practices: when 49% request truthfulness but define it differently, no single reward model can satisfy them. They connect this to broader concerns about flattening situated, contested preferences into universal models — a practice others have characterized as epistemic violence. The high persistence of hallucination rates in well-funded models, despite users' clear demand for accuracy, further suggests current methods fail to identify actual preferences.
- Only truthfulness is requested by a majority (49%), all other values fall below 25% across 75 countries.
- Even 'truthfulness' hides incompatible definitions: sourced claims, expert opinions, or unpopular views.
- Capabilities like human-like behavior and guardrails are controversial — some want them, others reject them outright.
Why It Matters
Current RLHF may be imposing a single value system on a globally diverse user base, undermining AI alignment's core goal.