Research & Papers

New Test Reveals AI Agents Struggle With Real Business Tasks

If AI can't follow your company's real procedures, it can't help your job.

Deep Dive

Standard operating procedures, or SOPs, are the step-by-step instructions companies use for routine work. Hospitals use them to register patients, banks to verify customers, and logistics teams to check hazardous shipments. These procedures make sure work is done consistently and safely. But when AI agents try to follow them, they often stumble—because real SOPs assume background knowledge and judgment that AI simply doesn't have.

To measure this, researchers created SOP-Bench, an open benchmark that tests AI agents on real, expert-written procedures from 12 business domains. It includes over 2,000 tasks, each paired with actual tools and known correct answers. Instead of just seeing if an AI can write a plausible text response, this test checks whether it can actually complete the procedure—using the right tools, in the right order, and catching the same unwritten details a human employee would know.

The results are sobering. Testing 11 top AI models showed that newer isn't always better, and giving the AI more tools often made its performance worse. No single model handled every procedure well. For example, a model great at coding tasks might fail at patient intake, because the procedure requires interpreting vague instructions like "verify insurance" in two different places—something that relies on real-world knowledge about two different verification methods.

The takeaway for businesses: don't assume a smart-sounding AI can handle your day-to-day operations. You need to test AI agents on your specific procedures before letting them work with customers, patients, or shipments. SOP-Bench provides a way to do that, and its open framework means other teams can extend it to their own sectors. This is a step toward AI that genuinely helps us at work, instead of quietly making mistakes.

Key Points
  • SOP-Bench tests AI agents on 2,000+ real business procedures from industries like healthcare, banking, and logistics.
  • Even the newest AI models made mistakes, and giving them extra tools sometimes made things worse.
  • Before deploying AI at work, companies need to test it on their own actual procedures—not just generic demos.

Why It Matters

If AI can't handle routine business steps reliably, companies risk costly errors in areas like healthcare, banking, and shipping.

📬 Get the top 10 AI stories daily