New study reveals deepfake benchmarks may be measuring the wrong thing
Simple probes rival specialized detectors on popular benchmarks, raising validity concerns
A new preprint by Pagon, Shen, Asnani, and Liu audits deepfake detection benchmarks using a deliberately simple diagnostic: a linear probe on frozen self-supervised representations. The authors find that across video, image, and audio modalities, these general-purpose probes closely approach the performance of bespoke detectors trained specifically for deepfake detection. This suggests benchmarks may be measuring general modality understanding (e.g., recognizing artifacts common to all generated media) rather than forensic capabilities tailored to deepfakes.
The work also shows that generator-level difficulty on benchmarks is partly explained by Fréchet geometry in the same representation space. The authors advocate for a benchmark-audit view: before interpreting high scores as evidence of forensic understanding, we must ask how much of the benchmark is already solved by general-purpose representations. If benchmarks fail to reflect real-world threat models, the complexity of engineered detectors may be solving the wrong problem entirely.
- Linear probes on frozen self-supervised representations match or approach bespoke deepfake detector performance across video, image, and audio benchmarks
- Generator-level difficulty is partially explained by Fréchet geometry in the representation space
- Results suggest benchmarks may reward general modality understanding, not true forensic deepfake detection
Why It Matters
High benchmark scores may not mean robust detection; real-world deepfake threats remain poorly measured.