Agent Frameworks

Nobody Agrees What an 'AI Agent' Is — Now There's a Scorecard

Companies keep selling you 'AI agents.' A new checklist could finally keep them honest.

Deep Dive

Two researchers, Mia Lassiter and Brinnae Bent, just published a survey paper on a surprisingly messy problem: nobody agrees on what an "AI agent" actually is. That matters because the word is everywhere right now. An agent (AI that can take actions for you, not just chat) might book your dinner reservation, file your expense report, or answer your customer emails. But when five companies all say their product is an agent, there's no standard ruler to compare them — so marketing language quietly does the work that evidence should.

The paper's answer is a five-part checklist. It asks whether a system can interact with its surroundings, learn and adapt, act without being told every step, work toward a goal, and stay coherent over time — meaning it remembers what it started and finishes the job. Think of it like food labels before regulators defined what "organic" meant: once there's a shared definition, shoppers can tell real claims from empty ones.

The authors also built the Agent Compendium, a free public website that organizes the tests and scoring methods researchers use to evaluate these systems. For you, the practical payoff is comparison shopping. If a bank, airline, or employer says it's deploying an AI agent, a common yardstick makes it easier to ask sharper questions: Can it act on its own? Does it remember context? What happens when it's wrong?

The catch: this is a map, not a verdict. The paper doesn't test any specific product, and no company is required to use these standards. It's a proposal from academics, and definitions in a fast-moving field tend to shift. Still, shared language usually arrives before shared rules — and this is one of the first serious attempts to build it.

Key Points
  • The word 'AI agent' has no agreed-upon definition, so company claims are hard to compare or verify
  • The authors propose a five-part test: can it interact, learn, act independently, pursue goals, and stay consistent?
  • They built a free public website, the Agent Compendium, collecting the methods used to grade these systems

Why It Matters

A shared scorecard helps you tell real AI helpers from hype before you trust one with your money or data.

📬 Get the top 10 AI stories daily