AI Safety

Adversarial psychometrics: a new way to measure superhuman AI intelligence

When AIs outsmart tests, let them quiz each other to find the true scale

Deep Dive

For decades, AI benchmarks relied on fixed, human-generated questions, but as systems approach or exceed human expertise, that measurement breaks down. A new academic paper by Elad Hazan, Kia Ghods, Jerry Han, Andrew Tu, and Rafael Moschopoulos proposes "adversarial psychometrics" to fill the gap. Drawing on Charles Spearman's 1904 concept of general intelligence (g) and Alan Turing's operational definition of intelligence, they argue that absolute performance on curated tasks becomes uninformative when a machine can outperform its examiners. Instead, they revisit Louis Leon Thurstone's 1927 Law of Comparative Judgement, the same principle behind Elo ratings in chess: pairwise comparisons can recover a latent capability scale without a fixed standard. This allows metrics to stretch infinitely as stronger participants enter the population.

The proposed protocol pits two AI agents against each other: one generates a challenge, the other attempts it; both are rewarded for separating their capabilities. Results aggregate into Bradley-Terry ratings, like Elo. The key innovation is avoiding external adjudication. Previous attempts, such as Token Games and MathDuels, restricted tasks to mechanically verifiable domains like logic puzzles and math, which excludes open-ended reasoning. Using another model as a judge simply shifts the problem—judges must be at least as capable as the systems they assess. Adversarial psychometrics instead rewards problem-validity and solvability implicitly, allowing the frontier of reasoning to be probed without a predefined rubric. This makes it a promising, scalable benchmark for the era of superhuman AI.

Key Points
  • Adversarial psychometrics lets AI models generate and solve each other's challenges, with rewards for separating capabilities, eliminating the need for human-built benchmarks.
  • Scales beyond superhuman intelligence through pairwise comparison (Thurstone's 1927 law, Elo rating), avoiding fixed standards that become uninformative.
  • Unlike verifiable-only protocols (Token Games, MathDuels), it handles open-ended reasoning without relying on an equally capable external judge.
  • Grounding in Spearman's general intelligence (g) and Turing's operational test offers a statistical, latent-scale measurement approach.

Why It Matters

As AI surpasses human benchmarks, adversarial psychometrics provides a scalable, judge-free way to rank models and guide development.

📬 Get the top 10 AI stories daily