Research & Papers

The AI That's Supposed to Read Your Dashboards Just Failed a New Test

⚡If you dream of AI doing your data analysis, this study says: not yet.

Deep Dive

Picture the dashboards you use at work: sales numbers, website traffic, budget charts. To answer one question you usually have to click a filter, then another, then hover over a chart, then remember what you saw two steps ago. Researchers want AI "agents" — software that can look at a screen and click things on your behalf — to do that clicking for you. The catch: nobody could tell exactly how badly today's AI performs that job.

A team led by Chuhan Zhang and Yingcai Wu built DashAct, a test designed to find out. Instead of only recording whether the AI got the final answer right, DashAct watches every step. It contains 357 real click-through sessions recorded and checked by humans, each with clear milestones — like "first filter by region, then compare two charts." The test then adds help in stages: it confirms what the AI has done so far, then offers hints about the next action, then finally points at the exact spot on screen to click. By measuring how much help the AI needs to recover, researchers can see where it actually breaks.

The results are sobering. Current models struggle even as researchers pile on support. The staged test exposed specific weak spots that a simple pass/fail score would have hidden — for instance, whether the AI forgets earlier steps, taps the wrong control, or simply cannot locate the chart it's supposed to read. In other words, the AI doesn't fail for one reason. It fails for several, and the biggest one varies by task.

Why should you care? Because the pitch for AI assistants is that they'll soon handle exactly this kind of routine screen work: pulling your weekly report, checking the numbers, flagging the anomaly. DashAct suggests that's further off than the marketing implies. For now, treat an AI assistant in a dashboard as a helpful pair of eyes, not a substitute analyst. The upside is that the researchers published where models break, giving developers a concrete map for fixing them — and giving you a clearer sense of when to trust the next version.

Key Points
  • A new test called DashAct checks not just whether AI finishes a dashboard task, but exactly where it gets stuck along the way.
  • It's built from 357 human-verified click-through sessions with milestones, so researchers can add help step by step and see what the AI truly needs.
  • Even with extra hints, today's models still struggle — meaning AI isn't ready to run your reports or pull your numbers unsupervised.

Why It Matters

Don't hand your weekly dashboard reports to an AI assistant yet — it may misread the numbers.

📬 Get the top 10 AI stories daily