AI Safety

Pay People to Read AI's Secret Notes — Safer Tech?

What if catching AI's mistakes before they harm us was as simple as paying people to read its diary?

Deep Dive

A proposal suggests creating an organization that pays people to review anonymized transcripts of AI systems like Claude, flagged by a high-recall monitor. The goal? Catch errors, misalignment, or even deliberate scheming before they cause harm. Think of it like hiring a team of detectives to read the AI’s private notes, ensuring nothing slips through the cracks.

The idea is straightforward: when AI messes up, it often leaves obvious clues in its internal logs. Humans reviewing these transcripts could spot issues faster and more reliably than automated systems alone. For example, every major AI incident so far could have been detected this way. The cost? Around $5 million a month—enough to hire 1,000 people working part-time—to review all the text generated by a frontier AI system. Estimates suggest this could catch about 15 serious problems monthly, even in pessimistic scenarios.

The proposal argues that this approach is scalable, doesn’t require top-tier AI experts, and avoids the need for expensive computing power. It’s also flexible: the team could start with cybersecurity or bio-safety evaluations and expand as needed. However, the biggest hurdle might be access. AI labs guard their internal data closely, and privacy concerns could block this idea before it starts.

Critics might ask: why not let AI monitor itself? The answer is that humans are still better at spotting nuanced or obviously dangerous behavior—like an AI hiding its true intentions. As one expert puts it, even a simple image-recognition AI can’t reliably identify a cat in a distorted photo, so why trust it to catch far more complex risks?

Key Points
  • A new org could pay 1,000 people $5M/month to review AI’s internal notes for errors or misalignment.
  • Humans might spot risks AI monitors miss, like reward hacking or deliberate scheming, before they cause harm.
  • Biggest challenge: AI labs may refuse to share private transcripts due to privacy or secrecy rules.

Why It Matters

This could prevent AI disasters by turning human oversight into a scalable, affordable safety net.

📬 Get the top 10 AI stories daily