Qwen 2.5 14B best mimics affordable-housing survey but fails on population structure
LLMs match averages but miss the social and spatial divisions of real-world opinions.
A new arXiv study (2607.27100) tested whether eight open-weight LLMs—including Qwen 2.5 14B and Phi-4 14B—could replicate the spatially anchored, identity-conditioned responses of 843 US survey participants on an affordable-housing proposal. The experiment measured how support changed as the project moved from 2 miles to 1/8 mile away, focusing on the gap between homeowners and renters. Qwen 2.5 14B was the only model to meet the prespecified ±0.20 equivalence criterion, producing a contrast of -0.242 versus the human -0.285. Phi-4 14B was directionally aligned but attenuated (-0.150), while other models showed weak, null, or reversed moderation.
Yet this aggregate match masked deep structural failures. Qwen 2.5 14B attenuated the Republican contrast and exaggerated the Independent one, yielding an RMSE of 0.613 across 27 party-by-tenure-by-item cells. Its median model-to-human variance ratio was just 0.099, and question order shifted the contrast by +0.367. Identity-cue removal and selective nonresponse changed which comparisons were estimable, and rationale-first responses differed from direct-choice ones in 20.6–35.3% of focal comparisons. The authors conclude that LLMs can approximate one average effect while failing to preserve the population structure, within-group heterogeneity, and measurement stability that generate it, urging urban planners to test spatial and social structure—not just averages—when using LLMs as low-cost proxies.
- Qwen 2.5 14B matched the human owner-renter difference (-0.242 vs -0.285) within ±0.20 equivalence criterion, outperforming Phi-4 14B (-0.150) and others.
- The model's RMSE across 27 party-by-tenure-by-item combinations was 0.613, median variance ratio 0.099, and question order shifted contrast by +0.367.
- Selective nonresponse and rationale-first responses changed 20.6-35.3% of comparisons, showing LLMs fail to preserve within-group heterogeneity and measurement stability.
Why It Matters
For urban planners: LLMs can't yet replace real surveys without distorting critical social and spatial opinion structures.