Developer Tools

New Scorecard Reveals Which Apps Are Actually Ready for AI Assistants

⚡Most popular software quietly blocks AI helpers — and that cost lands on you.

Deep Dive

AI agents — software that can click, type, and complete tasks on your behalf — are already here. But right now they mostly work by staring at screenshots of your screen, the same way a person would, guessing where to click. A new paper from Carnegie Mellon-style software engineering research argues that's the wrong setup. Instead, every feature you can reach by clicking should also be reachable by an agent through a clean, structured doorway, complete with safety notes.

To measure how far real software is from that, the author built two tools. The first, Agent-Callable Feature Coverage, gives a product a single readiness percentage: how much of what humans can do, agents can also do. The second scores each feature on three simple questions. Can an agent reach it? Can an agent find and understand it? Can an agent use it safely? Each answer gets 0 to 3 points, for a maximum of 9.

The results are sobering. Across 30 software systems in five categories, most scored moderately on raw access but poorly on documentation and safety governance. In a hands-on test of ten real systems, agents completed 56 of 58 tasks when instructions were available. Remove the documentation and that dropped to 50 of 58 — proof that even features agents can technically reach become unusable when nobody explains them. Features with no machine doorway at all were never solved, as expected.

The encouraging part: results held across three different AI models, including Google's open Gemma model running on a local machine — no expensive cloud bill required. For you, this is about whether the AI assistant you're paying for can actually finish your expense report, or just talk about it. The paper's quiet warning: software makers will need to rebuild their products with agents in mind, and that work hasn't started in earnest.

Key Points
  • AI agents mostly work by looking at screenshots and guessing where to click, which makes them slow and error-prone
  • In tests of 10 real systems, agents finished 56 of 58 tasks when instructions existed — but only 50 of 58 without them
  • The scoring system rates each feature out of 9 points for reachability, findability, and safe use

Why It Matters

Your AI assistant's real usefulness depends on whether software makers open proper doorways for it.

📬 Get the top 10 AI stories daily