Developer Tools

CMU researchers slash AI agent costs 90% with smarter harnesses

Small language models match frontier LLMs on business tasks at 4% the cost.

Deep Dive

A new paper from Carnegie Mellon University researchers tackles the high inference cost of frontier LLM agents for routine business tasks. The team shows that swapping a large model into a harness designed for a frontier LLM causes small language models (SLMs) to underperform. Their solution, automated harness adaptation, lifts task difficulty from the model into the harness by automatically generating tailored instructions, tools, and orchestration loops via a meta-agent.

The results are striking: across seven business-oriented agentic tasks and three SLM families (e.g., Llama 3, Mistral, Phi), optimized harnesses improved performance on 16 of 21 task-model pairs. Seven pairs fully closed the performance gap with frontier LLMs. The best SLM agent recovered 89.7% of LLM performance at just 4% of the cost—a 90%+ savings. The approach works best on repetitive workflows and with SLMs that have sufficient base capabilities.

This research has immediate practical implications. By making SLM agents competitive for the majority of routine business workflows, companies can deploy AI automation at a fraction of current costs. The automated harness optimizer essentially turns a weakness of small models—their inability to reason as deeply—into a strength by shifting the cognitive load onto structured, reusable scaffolding.

Key Points
  • Automated harness adaptation improves SLM agent performance on 16 out of 21 task-model pairs across Llama 3, Mistral, and Phi families.
  • Best SLM agent recovered 89.7% of frontier LLM performance while costing only 4% as much to run.
  • Approach works best on repetitive business workflows where harness templates can be reused across instances.

Why It Matters

Slashing agent costs by 90% makes AI automation economically viable for small businesses and routine enterprise workflows.

📬 Get the top 10 AI stories daily