Research & Papers

AI That Builds Other AI Just Got Cheaper and Faster

The software you use could soon be improved by AI, not human engineers.

Deep Dive

A team of researchers has built a system called AIBuildAI-2.5 that acts like an automatic AI engineer. Instead of a person writing code, testing designs, and training models by hand, a group of AI agents (AI that can take actions for you) does the whole loop: it proposes a design, runs it, checks the result, and tries the next idea. The goal is to make building AI something more people and smaller companies can afford to do.

The hard part has always been cost. Testing each AI design takes expensive computing time, so older systems could only try a few ideas and were often guessing in the dark. AIBuildAI-2.5 fixes this in three ways. First, an AI "judge" scores each candidate design on how promising, realistic, and well-grounded it looks, and a "selector" picks the best one to try next — like a coach choosing which play to run. Second, a scheduler decides when to start training jobs based on which computers are free, so hardware isn't sitting idle. Third, a router sends easy chores to cheaper AI models and saves the smartest, most costly model for the hardest steps. That combination cuts both the bill and the waiting time.

The results are strong. On MLE-Bench, a set of realistic machine-learning engineering challenges built from real data-science competitions, the system earned a "medal" on 73.3% of tasks — about three out of four — taking first place. It also beat a solid baseline on six open-ended AI research tasks.

The catch: this is still heavy-duty work that needs serious computing power and humans deciding what problem to solve in the first place. The test tasks are simulations of real engineering, not the messy, ambiguous projects people actually face. Still, the direction is clear — AI that improves AI could mean better tools, lower prices, and a real shift in what technical jobs look like.

Key Points
  • A new system lets AI agents design, train, and test other AI models with very little human help.
  • It placed first on MLE-Bench, earning medals on 73.3% of tasks — roughly three in four.
  • It saves money by routing easy steps to cheap AI models and reserving the priciest one for hard problems.

Why It Matters

Cheaper, faster AI building could mean better tools at lower prices — and more pressure on tech jobs.

📬 Get the top 10 AI stories daily