Research & Papers

RQGM: AI agents that co-evolve with their evaluators achieve 1.8x better paper acceptance

⚑Self-improving AI now evolves its own judges, beating static benchmarks by up to 1.86x.

Deep Dive

Self-improving AI agents have become SOTA on coding benchmarks, but they rely on fixed evaluators – a stationary utility that doesn't adapt as the agent improves. This ignores a fundamental evolutionary principle: species adapt as their environment changes. The Red Queen GΓΆdel Machine (RQGM) from researchers at Cambridge and Meta AI tackles this by making the evaluation criterion itself part of the improvement loop. The framework organizes search into epochs: within each epoch, the utility is fixed and standard self-improvement guarantees hold; between epochs, the utility can evolve, enabling adversarial objectives and dynamic utilities. On verifiable coding tasks, RQGM improves test pass rate over the prior SOTA while using 1.35x-1.72x fewer tokens by incorporating a complementary agent-as-a-judge code-review signal.

RQGM's most striking results come from scientific paper writing and grading. Co-evolved writer agents, training against a diverse panel of agent judges, achieve 1.78x-1.86x higher acceptance rates compared to agents trained with static evaluators. In grading, co-evolved graders reach 9% higher ground-truth accuracy. Perhaps most importantly, RQGM addresses a critical bias: the strongest baseline reviewer over-accepts AI-generated papers at up to 1.91x the human rate. By introducing an adversarial objective, RQGM discovers reviewers equally stringent on AI and human work, effectively correcting the bias. This work opens the door to truly recursive self-improvement where agents and their evaluators co-evolve, potentially leading to more robust and adaptable AI systems.

Key Points
  • RQGM uses 1.35x-1.72x fewer tokens than prior SOTA on coding benchmarks while improving pass rate.
  • Co-evolved scientific writer agents achieve 1.78x-1.86x higher acceptance rates under diverse judge panels.
  • RQGM corrects reviewer bias against AI-generated papers, reducing over-acceptance from 1.91x to parity with human work.

Why It Matters

By co-evolving agents and evaluators, RQGM unlocks adaptive, self-improving AI that goes beyond static benchmarks.

πŸ“¬ Get the top 10 AI stories daily