Research & Papers

New AI Called AREX-2 Grades Its Own Work — And Keeps Improving

⚡AI that spots and fixes its own mistakes could save you hours of redoing work.

Deep Dive

Researchers built AREX-2, an LLM agent designed to get better at a task by iteratively refining its solution at test time. The team trained it on long-horizon improvement trajectories synthesized from machine learning and algorithmic programming tasks — domains with verifiable feedback. Built on Qwen3.8-27B, AREX-2 scored 81.8 on MLE-bench Lite and 70.7 on Frontier-CS, and transferred to deep research with 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA — still improving as its budget of rounds grows. The authors conclude that long-horizon reflective data is an effective route toward self-improving agents.

Key Points
  • AREX-2 is an AI 'agent' — software that takes actions on its own — that reviews its own work and retries until the answer improves.
  • It scored 92.2 out of 100 on GAIA, a tough research test, after being trained on coding and data-science puzzles.
  • The more rounds it was allowed, the better it performed — but only in areas where correctness can be automatically checked.

Why It Matters

AI that catches its own errors means fewer wrong answers and less time double-checking its work.

📬 Get the top 10 AI stories daily