Research & Papers

New AI Trick Turns Months of Tedious Data Labeling Into Days

Pointing human experts only at where two AIs disagree saves months — and works better.

Deep Dive

Organizations increasingly need millions of documents sorted and tagged — customer reviews, medical notes, school transcripts. To do that, experts write a "codebook": a detailed rulebook saying what counts as what. This rulebook is what AI systems follow when they label text at scale. The problem is that writing a good rulebook is painfully slow. It can take experts months of reading, arguing, and rewriting before the AI gets it right.

A research team tested a shortcut. Let AI models start labeling the data, then compare two different AI models' answers. Wherever the two models agree, move on. Wherever they disagree, that's exactly where a human expert should look. They tried three ways for experts to weigh in: editing AI-written rule changes, simply answering questions about the disagreements, or labeling the tricky cases and explaining their reasoning.

They tested this on thousands of real tutoring-session transcripts. The winner was having experts label the borderline cases with explanations — that hit 64.9% accuracy against expert judgments, beating the traditional expert-revised rulebook at 57.8%. Just answering questions about disagreements also beat the old method, at 60.5%. In other words, a process that took months shrank to days, and the results got better, not worse.

The bigger lesson goes beyond rulebooks. AI is very good at the easy, obvious 90% of a task — and noticeably worse at the fuzzy, judgment-call 10%. Instead of asking humans to review everything, you can let AI flag its own uncertainty and spend scarce human attention only there. The catch: 65% accuracy is still far from perfect, so AI-labeled data can't be trusted blindly, especially in high-stakes settings like hiring, medicine, or grading.

Key Points
  • Experts only need to review the cases where two AI models give different answers — not everything the AI produces.
  • Letting experts label those tricky cases with explanations beat the traditional method: 64.9% accuracy versus 57.8%.
  • The approach shrank a months-long process to days, pointing to a new way of dividing work between people and AI.

Why It Matters

Cheaper, faster data sorting — and a preview of how humans and AI will split work.

📬 Get the top 10 AI stories daily