AI Safety

Cheap Specialist Judge Fails to Reduce Alignment Audit Costs

Lightweight Gemma 2-2B model used by agents but Sonnet API still eats 97-99% of costs.

Deep Dive

Researcher burnssa tested a cheap Gemma 2-2B toxicity-scorer as an audit tool for AuditBench’s investigator agents. The judge was used (~7–16 calls per audit) but barely reduced costs—Sonnet driver consumed 97–99% of expenses, and mandated tool use raised costs ~17%. Only helped in specific cases where the quirk type matched its training and the auditor struggled. Total experiment cost: $500+.

Key Points
  • Gemma 2-2B judge was used in every audit (7-16 calls per run) but only helped when quirk type matched training and Sonnet struggled.
  • Sonnet driver consumed 97-99% of audit costs; mandated tool use raised costs 17% and did not reduce total driver turns.
  • Total experiment cost exceeded $500, highlighting the expense of current alignment auditing methods.

Why It Matters

As AI models grow more capable and misalignment risks escalate, cheaper and more discerning audit tools are urgently needed.

📬 Get the top 10 AI stories daily