The AI Rankings Everyone Quotes Aren't Broken — Here's What They Hide
One number can't tell you which AI is actually best for your job.
Every week, a website called Artificial Analysis publishes a ranking of the world's best AI models. Recently, lots of people online started calling it "broken" and "bought off." This week, one frustrated fan wrote a long defense — and his points are worth knowing, because that ranking quietly shapes which AI tools companies buy, and which ones you end up using at work.
The defense is mostly about money and honesty. Artificial Analysis pays for its own testing instead of running ads, and it spent $13,129 to test just one model. Most of its tests come from public research papers anyone can read. It also publishes exactly how it blends ten different tests into one final score — unusually transparent for an industry where "just trust us" is the norm.
Then comes the part that actually matters to you. DeepSeek's new V4.1-Flash — a huge model, 552 billion parameters (think of parameters as the dials inside an AI; more dials means bigger, not automatically better) — got the same overall score, 40, as a much smaller rival, Qwen 3.8-Flash-Next. Same number, very different reality. DeepSeek beat even GPT-6 Astra at agentic tasks (AI that does things for you, like clicking through software and filing forms), but it was clearly worse at not making things up. A single score flattens all of that into one useless-looking digit.
The takeaway is simple. Before you pick an AI tool off a leaderboard, look at the specific test that matches what you need. If accuracy matters — say, for research or legal work — the "hallucination" score (how often the AI confidently invents facts) matters far more than the headline number. If you want AI to handle repetitive software chores for you, the agentic score is the one to check.
- Artificial Analysis, the site that ranks AI models, is being called 'broken' — but it funds its own testing and publishes exactly how it calculates scores.
- Two AI models both scored 40 overall, yet one was far better at doing tasks and the other far better at not making things up.
- Picking an AI off a single leaderboard number can lead you to the wrong tool — check the specific test that matches your need.
Why It Matters
The AI ranking you trust to pick tools hides big differences — you could choose the wrong one.