New AI alignment study: Pairwise comparisons fail when people hold conflicting values
Forced choices distort what users really want—allowing indecision cuts queries by 40%.
Flanigan and Si tackle a foundational assumption in AI alignment and participatory design: that asking people to repeatedly compare two decision rules (pairwise comparisons) can reliably reveal their true preferences. Their formal model introduces 'internal pluralism'—the idea that a single individual evaluates rules according to multiple authoritative priorities, such as fairness, efficiency, or equality. This internal conflict breaks the two core assumptions behind pairwise methods: that local comparisons suffice to capture global preferences, and that people can always answer decisively.
The authors identify two distinct failure modes. First, priorities like proportionality, egalitarianism, and equal treatment are inherently global—what they require in one case depends on outcomes in other cases, so local comparisons miss them entirely. Second, even when priorities are local, tension between strongly held values forces users into inconsistent choices when they must pick one option, leading to distorted behavioral data. Crucially, the paper shows that allowing people to say 'I can't decide' dramatically reduces the number of queries needed to learn preferences accurately. This opens the door to methods that directly elicit the underlying priorities, producing more faithful and interpretable models of what people actually value.
- Forced pairwise comparisons fail to capture inherently global priorities like proportionality or egalitarianism.
- Internal pluralism—holding multiple conflicting values—causes behavioral distortions in forced-choice settings.
- Allowing users to report indecision can reduce the number of queries needed for accurate preference learning.
Why It Matters
Improving how we learn human values is critical for building AI systems that truly align with diverse user preferences.