Praxa AI's Report Card Doesn't Say What It Seems, Audit Finds
If you pick AI tools based on performance numbers, this is a warning
A quiet research paper has taken apart the paperwork behind Praxa AI's internal test results, and the findings are a useful warning for anyone who shops for AI tools. "Agents" here means AI that can take actions for you — booking, filing, looking things up. The authors did not accuse anyone of cheating. They found numbers that added up correctly but secretly described something different from what their labels promised. One routing test listed 112 passes and 27 failures, yet no failures counted against the system, because known gaps were exempted from the pass-fail gate. So the "pass rate" was really telling you about the rules, not the performance.
The timing numbers are the clearest example. Across 8,843 records of AI tool attempts, 448 had no duration at all. Of the rest, 121 showed a duration of 2,147,483,647 milliseconds — the largest number a common computer format can hold — and were tagged as "abandoned client." That made the headline slowest-response figure come out at roughly 24 days, compared with about 38 seconds for calls the server actually finished. It looks like a dramatic slowdown, but it is really two different groups of records being mixed together.
There is a similar gap in a cost-saving experiment. Praxa AI reported that a technique for shrinking old conversation text — a way to cut the bill for running AI — reduced input by 94.39%. Counted across the entire event rather than just one follow-up step, the real reduction was 46.54%. That is still a real saving, just less than half the advertised figure.
Here is what it means for you. If you or your employer choose AI products based on vendor dashboards — speed, accuracy, pass rates — those numbers may be measuring company policy choices or computer quirks rather than daily reality. The honest catch: this is one retrospective audit of one company's public files. The authors did not rerun the AI itself and make no claim it is worse than its rivals. The lesson is simpler and broader: always ask how a number was made.
- An audit of Praxa AI's public test files found metrics that were mathematically correct but described something different from their labels.
- One reported slowest response time read as about 24 days, because 121 timing records used a computer placeholder meaning "gave up" instead of a real duration.
- A claimed 94% cut in processing cost was closer to 47% once the whole task was counted, not just one step.
Why It Matters
Ask how AI speed and accuracy numbers were measured — a label may not match the math behind it.