LLM agent co-design creates overtrust trap, study finds
Users think their AI twin is spot-on — until independent tests prove otherwise.
A new paper by researcher Michael J. Fell, published on arXiv, examines a troubling paradox in the design of LLM-based preference agents. These agents are increasingly used to simulate human opinions in research and product development. Co-designing them with the people they represent seems like a natural fix for issues of misrepresentation and exclusion. However, Fell’s qualitative study with 12 participants in the household energy domain reveals that the act of participation can actually worsen trust calibration.
Participants completed a background survey, a co-design interview, and a validation survey. They engaged enthusiastically and generally felt their agents represented them well. But when the agents were tested independently, alignment was poor. The agents produced responses that were significantly more homogeneous, more decisive, and more abstract than the human sample. Fell argues that participation and process transparency create an “overtrust engine” — users trust the agent because they helped build it, even when it systematically misrepresents them. The paper warns that deploying such agents at scale could embed structural biases while giving users a false sense of control. Fell treats alignment not as a fixed state but as an ongoing enacted process that requires constant external validation, not just user feedback.
- 12 participants co-designed LLM preference agents for household energy decisions through a survey, interview, and validation loop.
- Independent validation showed agents were markedly more homogeneous and decisive than the human sample they were meant to represent.
- Author introduces 'overtrust engine' concept: participation masks systematic misalignment, leading to misplaced confidence in agent accuracy.
Why It Matters
As LLM agents simulate human preferences at scale, this study warns participation alone can't guarantee trustworthy alignment.