AI Test Scores Are Misleading: Many 'Impossible' Tasks Were Just Broken
If the tests are rigged or broken, the AI story you're sold is wrong.
Here's the setup. AI companies love to brag that their models pass tough tests. One popular test suite, called Terminal-Bench, gives AI "agents" (software that can actually take actions on a computer, not just chat) a list of chores to finish in a text-based terminal window. The researchers pulled a frozen record of this testing: 1,081 submitted tasks, 639 that were scored, 28,801 individual attempts, and $105,933 in real money spent running the AI. Then they asked a blunt question: when every single model fails a task, what does that actually prove?
The answer, mostly, is "not much." They looked at the 125 tasks where no AI ever passed honestly, checking the answer keys, rerunning the official solutions, trying deliberately empty attempts, and hunting for loopholes. Only 78 held up as genuinely unsolved. Fourteen had broken or faulty answer keys. Eight failed because of computer problems, not AI limits. Four could be beaten by tricking the grading system. Twenty-one simply couldn't be verified either way.
So why should you care? Because these numbers become headlines, funding decisions, product prices, and even arguments about which jobs are safe. If a company says "no AI can do this yet" — or claims the opposite — you deserve to know whether the test was actually fair. Bad benchmarks quietly distort what you're told about AI's real abilities in the workplace.
The honest catch: even the 78 "certified unsolved" tasks aren't proven to be truly impossible. That label only means the official solution worked, no computer glitches dominated, nobody cheated, and every tested AI failed. The authors say benchmark makers should publish the evidence behind their failing scores before anyone treats them as proof of anything.
- Of 125 AI test tasks where every model failed, only 78 survived a careful audit.
- Problems included 14 broken answer keys, 8 computer failures, and 4 tasks you could cheat on.
- One test run alone cost $105,933 — so bad benchmarks waste real money and mislead buyers.
Why It Matters
Misleading AI test scores shape what you're sold, what gets funded, and which jobs people fear losing.