Developer Tools

Your Company's AI Analyst May Be Confidently Wrong — New Test Finds Out

The smarter AI answered more questions but quietly changed its story 41 times.

Deep Dive

Companies are starting to use AI agents (software that can take actions on its own) to answer business questions. You ask, "Why did sales drop in Europe?" and the AI picks the data sources, runs the numbers, and writes an answer that might land in a board deck. A new research paper argues that judging these tools by their final answer alone is dangerous, because a good-looking answer can quietly come from the wrong data.

The team built a testing method that grades the whole journey, not just the destination. They check three things: whether the AI understood the question, whether it ran the steps correctly, and whether it gives the same answer twice. They tested an internal analytics agent at a large online marketplace using 50 real business questions, two AI setups, and three repeat runs each — 300 runs in total.

The results are a warning. The more capable AI refused to answer only 0% of the time, down from 73%, and gave real-data answers 73% of the time, up from 21%. That sounds great. But it also ran out of its allowed steps on 16% of runs, blew past its data-exploration limit on 77% of runs, and changed how it read the same data tables on 41 of 50 questions. It was more helpful and less reliable at the same time.

The lesson for anyone buying or using these tools: ask how they were tested. A demo that answers one question well proves almost nothing. You need repeated runs, visible step-by-step records, and honest reporting of how often the AI gives up or contradicts itself — especially for anything touching money, hiring, or strategy. The paper's authors say these scores should guide decisions, not act as a simple pass-or-fail gate.

Key Points
  • The paper grades AI business assistants on their full process — question understanding, execution, and consistency — not just whether the final answer looks right.
  • In testing, the stronger AI model answered far more questions (73% vs 21%) but ran out of allowed steps in 16% of runs and overran its data-search limit in 77% of them.
  • It changed its interpretation of the same data tables on 41 out of 50 questions — a sign that confident-sounding AI answers can shift without warning.

Why It Matters

If your company trusts AI for reports or forecasts, ask how it was tested — a good answer can hide bad data.

📬 Get the top 10 AI stories daily