AI Safety

Alex Mallen questions judgment prediction for AI conceptual benchmarks

Expert judgments are noisy, and models might just learn the judge's prior views.

Deep Dive

In a recent post on the AI Alignment Forum, Alex Mallen explores whether judgment prediction tasks—where an AI predicts a specified expert's opinion on subjective questions—could serve as a benchmark for conceptual reasoning capabilities. Mallen argues that many important conceptual tasks (e.g., estimating the probability of misaligned AI takeover) involve hard-to-resolve disagreements, making traditional objective benchmarks unsuitable. The proposed methodology would involve paying experts to answer questions under controlled conditions (e.g., with 10–60 minutes of thought, tools, or AI assistance), then measuring how closely LLMs can predict those judgments. The goal is to track AI progress on disagreement-laden tasks without conflating capability with alignment of tastes or priors.

However, Mallen highlights several critical downsides. Human judgments are inherently noisy, and because people have memory, you cannot reset and resample them to measure that noise precisely—unlike in objective datasets where multiple judges can cross-check. For instance, early experiments with "Mythos preview" showed LLMs nearly matching Ryan Greenblatt's own accuracy on multiple-choice questions about his unpublished views, suggesting models may simply be learning the judge's prior stances rather than performing deeper reasoning. Data leakage is another major concern: if a model trains on statements like "Caspar finds MIRI’s work interesting," it could mechanically predict Caspar's subsequent judgments without understanding the reasoning. Score improvements driven purely by later knowledge cutoffs (e.g., learning the judge's evolving public views) would further invalidate the benchmark as a measure of conceptual capability. Ultimately, while judgment prediction might point to deficits in AI reasoning, noise makes it nearly impossible to confidently assert that a higher score reflects genuine improvement rather than memorization or luck.

Key Points
  • Judgment prediction tasks aim to benchmark AI on subjective conceptual questions by having models predict a specific expert's view, isolating capability from taste disagreements.
  • Early tests (e.g., Mythos preview) show LLMs can match expert self-prediction on multiple-choice questions about unpublished views, indicating potential data leakage issues.
  • Mallen warns that human judge noise, memory effects, and knowledge cutoff artifacts make it hard to verify that score gains reflect genuine reasoning improvements rather than mere memorization.

Why It Matters

For AI alignment researchers, reliable benchmarks are critical—this proposal highlights fundamental validity challenges in measuring conceptual reasoning on subjective topics.

📬 Get the top 10 AI stories daily