Developer Tools

New Tool Pinpoints Which Step an AI Assistant Got Wrong

This could make AI helpers far more reliable — and cheaper to fix.

Deep Dive

Imagine asking an AI assistant to build you a sales report over five back-and-forth messages. In step two, it accidentally pulls "profit" numbers instead of "revenue." It doesn't crash. It doesn't warn you. It just keeps going, and every answer after that is quietly wrong. You get a polished report at the end that's built on the wrong numbers.

That's the problem a new evaluation method called AEM (short for Agent Evaluation Metric) is trying to solve. Today's grading tools only look at the final result: did the assistant finish the task or not? Most give you one number — something like "70 percent success" — with no clue whether the failure was a wrong fact, a missing detail, or a bad decision about which tool to use. That's like a teacher handing back a test marked "C" with no notes on which questions you missed.

AEM instead breaks quality into named pieces you can measure one at a time. In this first version, it checks two things: truthfulness (are the numbers and facts right?) and completeness (did anything get left out?). It scores every single turn of the conversation, then traces backward to find the exact step where things went wrong — and, crucially, separates that root cause from the later steps that only inherited the problem. The researchers say the same approach can later be extended to other concerns, like safety or whether the assistant remembered your instructions.

Why should you care? AI agents — software that takes actions on your behalf, not just chats — are being rolled out for customer service, scheduling, shopping, and workplace admin. Their weakness isn't usually one big blunder; it's small early errors snowballing silently. A tool that shows exactly which step failed makes these systems faster and cheaper to fix, and it gives companies a reason to be honest about where their AI still breaks. For anyone relying on an AI assistant to get real work done, that's the difference between trusting it and checking everything twice.

Key Points
  • One small error early in an AI conversation can quietly corrupt every answer after it — and today's grading tools can't tell you where it started.
  • The new AEM method scores each step separately, checking two things: whether facts are right and whether anything was left out.
  • It's aimed at AI 'agents' — software that books, schedules, and does tasks for you — where hidden mistakes cost real time and money.

Why It Matters

More reliable AI helpers means less time double-checking their work and fewer costly mistakes on tasks you delegate.

📬 Get the top 10 AI stories daily