New Paper Exposes How LLMs Can Be Used to Manipulate Truth
Researchers reveal a framework for detecting and auditing AI-driven manipulation in public discourse.
A team of researchers from the University of Toronto has published a sweeping 50-page paper titled "Adversarial Social Epistemology for Assemblies of Humans and Large Language Models." The authors, Mihnea C. Moldoveanu and Joel A.C. Baum, argue that current models for understanding misinformation—like epistemic bubbles, echo chambers, and simple diffusion—fail to capture how agents actively exploit trust structures in complex communicative landscapes. They propose a new framework called Adversarial Social Epistemology (ASE) that focuses on how both humans and LLMs distort, omit, fabricate, or strategically underspecify information for private, reputational, or material gain. The paper draws on inferentialist semantics and epistemic networks to map out how assertions are scaffolded by chains of testimony, certification, and tacit trust.
The core innovation lies in describing how communicative agents subvert the very commitments and entitlements that make scaffolded assertions trustworthy. The researchers outline specific mechanisms that erode auditability of inferential chains—for example, by exploiting automated LLM outputs or seeding plausible but false narratives into institutional certifications. They also provide machinery for auditing these chains and redressing trust breaches. While the paper is theoretical, its direct relevance to today's AI-driven information environment is clear. As LLMs like GPT-5 and Claude are increasingly used to generate news, reports, and social media content, the ASE framework offers a systematic way to detect manipulation that goes beyond simple fact-checking. The authors suggest that platforms and regulators could use these ideas to build more robust trust architectures.
- Introduces Adversarial Social Epistemology (ASE) to model truth distortion in human-LLM communication chains
- Goes beyond echo chambers and misinformation to focus on how agents exploit trust in inferential chains
- Provides specific auditing machinery to detect and redress breaches of trust in public assertions
Why It Matters
As LLMs become embedded in public discourse, ASE gives us a systematic way to detect sophisticated manipulation that fact-checking alone misses.