Research & Papers

New BiBiR framework exposes LLM disinformation detector flaws

LLM-generated disinformation evades detection 95% of the time with adversarial rewrites

Deep Dive

Researchers from the University of Sheffield (led by Carolina Scarton) have developed the Build it, Break it, Repeat (BiBiR) framework to expose critical weaknesses in LLM-based disinformation detection systems. The team iteratively pitted "builders" (models designed to detect manipulated content) against "breakers" (adversarial techniques to evade detection) in five rounds of testing.

The strongest adversarial breakers achieved a 95% label flip rate (LFR) by combining back-translation and LLM persona-based rewriting while preserving the original post's meaning. Meanwhile, the top-performing builder—a triplet contrastive model with dynamic anchor switching (DASS) architecture—achieved 72.68% average accuracy, outperforming a fine-tuned e5-small-LoRA baseline by 15 percentage points. The findings underscore that static benchmarks fail to capture real-world adversarial evasion, and iterative stress-testing is essential for building robust detection systems.

Key Points
  • BiBiR framework tests detectors by iteratively building adversarial attacks and improving defenses, revealing 95% label flip rates with preserved semantics
  • University of Sheffield's DASS-based detector achieved 72.68% accuracy, beating the e5-small-LoRA baseline by 15 points
  • Static benchmarks underperform against dynamic adversarial attacks, highlighting the need for iterative robustness testing

Why It Matters

This work forces a critical rethink of how we evaluate AI disinformation detectors under real-world adversarial conditions.

📬 Get the top 10 AI stories daily