Research & Papers

Researchers show tuned LLMs can strategically amplify social network dissensus

Reinforcement learning-tuned LLM approaches theoretical disruption limits on real-world networks

Deep Dive

In a new arXiv paper, computer scientists Erica Coppolillo and Giuseppe Manco demonstrate that large language models can be deliberately fine-tuned to amplify social discord by injecting targeted opinion content into networks. The researchers built on the Friedkin-Johnsen (FJ) model, a classic framework for how individuals update opinions by balancing their own beliefs with neighbors'. They showed that standard FJ variants are surprisingly resistant to perturbation—random or naive content injection barely moves equilibrium dissensus. However, by extending the model to allow shifts in an individual's inherent opinion, they unlocked valid graph structures where disruption at equilibrium exceeds the initial state, making strategic attack vectors viable.

To operationalize this, the team designed a reinforcement learning (RL) loop that fine-tunes an LLM to generate 'disruption-oriented text'—messages engineered to maximize dissensus across the modeled population. Experiments on both synthetic and real-world social graphs confirmed that the RL-tuned LLM consistently approaches the theoretical ceiling of disruption predicted by the extended FJ model. Notably, the method works without modifying network topology; it only changes what content certain users see or believe. The authors explicitly frame these findings as a warning: they expose vulnerabilities in current moderation systems and highlight how generative models could be weaponized for adversarial information campaigns. They also call for stronger regulation of generative model outputs, especially where fine-tuning is accessible. The code has been released publicly, making the method reproducible—and the implications urgent for social platforms, policy makers, and AI safety researchers.

Key Points
  • Standard Friedkin-Johnsen models resist perturbation, but extending the model to shift inherent opinions enables network disruption at equilibrium
  • A reinforcement learning framework fine-tunes an LLM to generate text that approaches theoretical dissensus limits on real-world graphs
  • The work highlights concrete risks for content moderation and adversarial campaigns, with code publicly released for reproducibility

Why It Matters

This shows LLMs can be weaponized to destabilize online discourse, demanding proactive moderation and AI regulation.

📬 Get the top 10 AI stories daily