Study: AI Contract Readers Flag Clauses That Aren't There
If AI reviews your contracts, it may invent problems — or miss real ones.
A new paper takes a hard look at something most of us never see: what happens when an AI reads a contract and tells you what's inside it. The dataset used was CUAD, a well-known public collection of 510 contracts that researchers use to test these systems. The author re-ran a fixed, publicly available AI model and measured two things people usually ignore — how often it stays quiet when it should, and how much evidence exists to trust each answer.
The headline number is uncomfortable. The system was supposed to pull out key clauses, and it did find 82.2% of the ones human lawyers had marked. But it also spat out 30,464 candidate phrases, and 19,562 of those came from questions where the contract had no answer at all. For those no-answer questions, the AI gave an answer 45.4% of the time. In plain terms: if you asked it to check whether a contract has a certain clause, and it didn't, the AI would often say something anyway. A lawyer still has to read every one of those to find out it's a false alarm.
There's a second problem, and it's about trust in the scoreboard itself. Of the 41 clause categories tested, only 9 had at least 30 contracts with the clause and 30 without. Change that threshold slightly and the number swings to 18 — or down to two. Two categories had zero examples of "clause missing," meaning you literally cannot calculate how often the AI gets them wrong. A single accuracy percentage hides all of that.
The honest limits: this is one model, one scoring setup, one dataset — not a verdict on every AI contract tool you might buy. The author also flags that confirming results is hard because the model may have seen the test data during training. But the lesson travels well: when AI reviews your paperwork, ask not just what it found, but what it claimed to find that wasn't there.
- The AI answered 45% of questions where the correct response was 'no such clause exists' — creating false alarms a human must then clear.
- It still caught 82.2% of real clauses, but produced 30,464 candidate phrases, nearly two-thirds of them from answer-less questions.
- Only 9 of 41 clause types had enough examples to judge reliably, so a single 'accuracy score' can hide weak spots.
Why It Matters
If your company uses AI to review contracts, budget time for false alarms — and don't trust one accuracy number.