Research & Papers

Humans and AI fail to detect synthetic legal evidence in new study

Neither humans nor frontier MLLMs reliably spot AI-generated legal photos...

Deep Dive

A new study from researchers at the University of Montreal, published on arXiv, reveals that both humans and cutting-edge multimodal AI models struggle to identify AI-generated synthetic evidence in legal contexts. The team built SLED-1400, a dataset of 200 authentic evidence photos (e.g., damaged property, accident scenes) paired with 1,200 AI-generated fakes from six top text-to-image systems including Gemini-3-Pro-Image, Flux-2-Max, GPT-5.1, and others. In a web experiment with 136 lay participants, human accuracy hovered at 64.8% overall and fell to 48.5% and 51.0% on the two strongest generators – effectively a coin flip.

Four frontier multimodal LLMs (GPT-5.1, Gemini-3-Pro, Gemini-3-Flash, Qwen3-VL-235B) were tested under identical conditions. They scored 100% specificity (never falsely flagging an authentic image as fake) but abysmal sensitivity: average detection rates for the hardest generator (Gemini-3-Pro-Image) were just 5.9%. MLLMs also showed strongly correlated error patterns while humans errors were uncorrelated with AI errors. The authors argue visual evidence should be treated as 'inherently contestable' and recommend a multi-layered approach: trained human review, MLLM screening, and mandatory provenance metadata like C2PA Content Credentials.

Key Points
  • Humans achieved only 64.8% accuracy overall, dropping to ~50% on top generators like Gemini-3-Pro-Image and Flux-2-Max.
  • Frontier MLLMs (GPT-5.1, Gemini-3-Pro, etc.) had 100% specificity but missed 94.1% of fakes from the hardest generator.
  • Researchers recommend combining human review, MLLM screening, and C2PA provenance infrastructure for legal evidence.

Why It Matters

As AI-generated visuals flood courtrooms, neither humans nor current AI alone can reliably authenticate legal evidence.

📬 Get the top 10 AI stories daily