Research & Papers

HOTE framework trains 8B model to outperform larger deep research AIs

By evolving three agents jointly, HOTE's 8B model beats 32B models

Deep Dive

A new paper from researchers Piao et al. introduces Hybrid Open-Ended Tri-Evolution (HOTE), a framework that bridges deep research and agent evolution. Deep research allows AI to autonomously retrieve and synthesize information in open-ended environments, but is limited by static parametric capabilities. Agent evolution, on the other hand, improves models through interaction but has only been verified on tasks with standard answers. HOTE unifies these by evolving three specialized modules—a proposer, solver, and judge—using hybrid-mode reinforcement learning on web-scale knowledge. This tri-evolution allows each component to improve collaboratively, moving toward fully autonomous research agents.

Extensive tests on three long-form deep research benchmarks show that HOTE's 8B model surpasses all static open-source models ranging from 8B to 32B parameters, as well as models trained with prior state-of-the-art deep research methods—and it does so with less training time. The paper confirms that all three evolved modules are essential; removing any one degrades performance. This work demonstrates that small, efficiently trained models can rival much larger systems for complex research tasks, making autonomous research agents more accessible and practical.

Key Points
  • HOTE co-evolves three agents (proposer, solver, judge) via hybrid-mode reinforcement learning on web-scale knowledge.
  • An 8B parameter model trained with HOTE outperforms static open-source models up to 32B and prior SOTA deep research methods.
  • The framework achieves better results with lower time overhead, and all three modules are indispensable for peak performance.

Why It Matters

Efficient, autonomously evolving research agents could accelerate scientific discovery without requiring massive models.

📬 Get the top 10 AI stories daily