Study: GPT-5.2 best matches human sympathy in news framing
GPT-5.2 hits 0.789 alignment with human empathy; Mistral Large scores 0.4
A new arXiv study from researchers Haran Shani-Narkiss, Michael Fire, and Oren Tsur examines whether LLMs grasp the emotional framing of news headlines the same way humans do. In a large-scale experiment, 3,011 demographically representative UK participants recruited via YouGov and seven LLMs—including GPT-5.2, Mistral Large 2512, and others—were asked whether headlines from political and geopolitical conflicts evoked sympathy for a specified side. The result: AI-human agreement varies dramatically, from a very high correlation of 0.789 for GPT-5.2 down to a medium 0.4 for Mistral Large 2512.
The study goes beyond aggregate accuracy to probe differential alignment—how AI judgments correlate with human subgroups. Even the best models showed statistically significant differences across age, gender, education level, prior geopolitical knowledge, and participants' predispositions regarding the conflict. While leading models appear broadly aligned with overall human perception, the authors argue that aggregate performance masks uneven alignment by demographic and cultural norms. This work, with its robust design and large dataset, provides the most comprehensive evaluation yet of LLM comprehension of news framing, underscoring the need for ethical AI that accounts for heterogeneous human sensitivities.
- GPT-5.2 achieved 0.789 correlation with human sympathy judgments; Mistral Large 2512 only 0.4
- Study used 3,011 UK adults via YouGov and 7 LLMs on political conflict headlines
- Even top models show significant alignment differences across age, gender, education, and political views
Why It Matters
High aggregate AI-human alignment doesn't guarantee fairness—differential alignment gaps must shape ethical AI design.