AI Safety

AI audit reports can sound plausible while citing irrelevant evidence, study finds

600 reports, 60 policy cases, 5 human reviewers reveal a hidden governance risk.

Deep Dive

As AI governance shifts from benchmark scores to auditable oversight, a central question is how reviewers can trust that an LLM-generated audit report is truly evidence-backed. Seunghyun Yoo's paper (arXiv:2607.21462) introduces a controlled evaluation framework that holds the policy passage, rubric, and auditor model fixed while changing only the evidence interface. Across 60 AGORA policy cases, the team generated 600 structured reports under 10 evidence conditions, including passage-based evidence, internal model evidence, a hybrid package, and a shuffled control that preserved evidence format but broke case relevance. Five human reviewers evaluated the primary interfaces for correctness, passage grounding, diagnostic usefulness, and evidence misuse.

The results reveal a nuanced picture: internal evidence changes how reports cite and reason about evidence, but more internal citations do not by themselves make a report more valid. A white-box diagnostic explains the failure mode—causal localization is narrow while reports readily reuse broader readable labels and token directions. The hybrid interface was rated the most useful on average. Critically, the shuffled control exposed a key governance risk: reports can sound substantively plausible while citing irrelevant internal evidence. This study reframes internal model access as an evidence design problem for audit workflows, rather than as a guarantee of transparency.

Key Points
  • 600 structured reports generated from 60 AGORA policy cases across 10 evidence conditions.
  • Hybrid evidence interface (passage + internal) rated most useful by human reviewers.
  • Shuffled control showed LLM audit reports can cite irrelevant evidence while sounding plausible.

Why It Matters

As AI governance mandates audits, this study shows evidence design—not just access—determines report trustworthiness.

📬 Get the top 10 AI stories daily