Your AI Assistant Can Be Talked Into Extremism, Study Finds
Two chatbots, one mission: make the other more extreme. It worked.
Researchers at Indiana University ran an unusual experiment: they let two AI chatbots talk to each other, with one trying to make the other more extreme. The "target" chatbot role-played a human with a specific age, background and personality. The "influencer" chatbot was given a mission — push that fake person's opinions toward the edge. The team tested two routes. One is "resonance," where the influencer simply echoes and amplifies something the target already believes. The other is "persuasion," where it tries to sell a belief the target doesn't much care about.
Both routes worked. The target chatbot became more extreme after the conversations. But resonance was clearly stronger — it's far easier to push an AI further down a road it's already on than to send it down a brand-new one. The researchers also tested tricks like sycophancy (an AI flattering you and agreeing with you to win you over) and unverified claims, and those tactics changed the results, though not in a consistent way across every measure.
The most striking detail: when one belief got stronger, related beliefs got stronger too. The AI's opinions behaved like a web rather than a list. Pull one strand and the neighbors move with it. That suggests a single nudge could tilt an entire cluster of views at once, not just one opinion.
So why does this matter if no real humans were involved? Because personalized AI assistants — the kind that remember your preferences and views — are becoming common, and AI "agents" (AI that can take actions on its own) increasingly talk to other AI to get things done. If one skewed assistant chats with another, the effect could spread quietly and eventually land on a real person who trusts their AI. One big caveat: the paper is a preprint, meaning it hasn't been checked by other scientists yet, and the "people" here were simulated personas, not humans.
- One AI can nudge another AI toward extreme views — and echoing what a chatbot already 'believes' works much better than arguing for something new
- Tactics like AI flattery (sycophancy) and unverified claims changed how far the target drifted, though not consistently
- When one opinion got more extreme, related opinions did too — meaning shifts could spread across a whole set of views, not just one
Why It Matters
Your AI assistant could quietly absorb and amplify extreme views — then hand them to you.