Developer Tools

Amazon Now Grades Your AI Assistant's Work, Step By Step

Companies can finally check whether their AI actually followed the rules

Deep Dive

Amazon Web Services has added new testing tools to its AI platforms (Strands Evals and Bedrock AgentCore) that grade how well AI assistants do specialized jobs. The idea centers on "skills" — think of them as recipe cards. Instead of stuffing every company rule into one giant instruction sheet, you write separate cards, each teaching the AI one task: how to redact a legal contract, how to reconcile an invoice, how to follow a team's coding habits. The AI loads only the card it needs. That keeps things easy to update, and it means no retraining of the underlying model is required.

The problem: AI can fail in two sneaky ways. It might grab the wrong recipe card entirely — answering a question about vacation time using the employee benefits policy. Or it might pick the right card, then skip half the steps. Both failures produce an answer that sounds fluent and confident but is actually wrong. Standard quality checks miss this, because they only look at the final wording, not the process behind it.

Amazon's new graders fix that. One checks whether each skill the AI picked was actually appropriate — a simple yes or no. Another rates, on a five-level scale, how completely the AI followed each prescribed step, with evidence. A third is a quick automated check that the right skill loaded at all. Crucially, these work from a recording of what the AI did, so companies can review past runs rather than redoing the work.

For everyday people, the payoff is fewer confidently wrong answers from chatbots handling refunds, billing disputes, or paperwork. The catch: this is a tool for developers. It only helps if the companies behind those assistants actually use it.

Key Points
  • AI 'skills' are like recipe cards — separate instruction files for each task, loaded only when needed
  • Amazon's new tests catch two sneaky failures: the AI grabbing the wrong recipe, or skipping steps in the right one
  • It grades recorded runs, so companies can audit past AI behavior without repeating the work

Why It Matters

Fewer confidently wrong answers from AI chatbots handling your bills, refunds, and contracts.

📬 Get the top 10 AI stories daily