AI Safety

Anthropic Test Shows AI Still Can't Judge Good AI Safety Research

AI can't yet tell which research makes AI safer — humans still needed.

Deep Dive

Anthropic researchers built a test called TASTE to see if AI models can judge the quality of AI safety research proposals. The idea is simple: give the AI pairs of research proposals and ask which one is more important and worth pursuing. Then compare the AI's choices with the preferences of experienced human researchers. The result? Humans agreed with each other 77% of the time, but the best AI model only matched human preferences 60% of the time. In other words, AI is still clearly worse than people at judging what research actually matters for safety.

Why does this matter? As AI becomes more capable, some researchers believe we might need AI to help do safety research automatically — because AI development could outpace human scientists. But if AI can't even reliably tell which safety research is more promising, then fully automated safety research isn't ready yet. Judging proposals well is a high-leverage skill: picking the wrong research direction wastes time and money, and in safety-critical areas, that could be dangerous.

Getting this benchmark right wasn't easy. The researchers used two key tricks to make human labels reliable. First, they let researchers discuss disagreements before changing their scores. Second, they only kept labels where researchers were strongly confident. That gave them 92 high-quality comparison pairs. Even so, the fuzzy nature of research judgment means 77% human agreement is about as good as it gets — there's no single "right" answer in many cases.

The big takeaway: for tasks that are hard to verify, like evaluating research proposals, AI still falls short of human judgment. The authors see this as a step toward building better benchmarks, not a verdict on AI's potential. But right now, if you want to know which AI safety research deserves funding, asking a human expert is still your best bet.

Key Points
  • An Anthropic test called TASTE found AI models are only 60% as good as human experts at judging AI safety research proposals.
  • Human researchers agreed with each other 77% of the time, showing research evaluation has no single 'right' answer.
  • The study used only 92 comparisons and filtered for confident human ratings to get reliable results.

Why It Matters

AI can't yet be trusted to decide which safety research matters — humans still must guide it.

📬 Get the top 10 AI stories daily