Agent Frameworks

New Human-on-the-Bridge paradigm scales AI agent evaluation with upstream expert curation

23,500 agent turns tested across finance, healthcare, and code generation...

Deep Dive

Current methods for evaluating AI agents — static benchmarks, human-in-the-loop review, LLM-as-judge, red teaming, and trace auditing — each provide only fragmented signals and struggle to scale. In a new arXiv paper, author Fouad Bousetouane proposes Human-on-the-Bridge (HOB), a paradigm shift that encodes human expertise upstream, before any test execution. Experts build a reusable evaluation intelligence bundle containing domain context, adversarial Red-Team Traps, Juror Personas (diverse scoring perspectives), detailed guidelines, audit rules, and fallback policies. The ProofAgent Harness then repeatedly executes this curated intelligence across multi-turn adversarial evaluations, capturing full traces, applying multi-juror scoring, and generating evidence-linked reports. This design allows smaller evaluator LLMs to effectively challenge agents built on frontier backbones.

HOB was evaluated in symmetric and cost-efficient asymmetric settings across finance, healthcare, and code generation tasks — totaling 23,500 agent turns. The results revealed failures that static benchmarks and single-evaluator scoring routinely miss: phantom tool-call claims, missing mandatory tool calls, policy drift during long interactions, subtle manipulation paths, and safe but non-resolving refusals. By moving human judgment upfront and reusing it, HOB amplifies evaluation quality without requiring equally large evaluator models for every test run. The approach promises to make rigorous, human-curated agent evaluation both scalable and repeatable for real-world deployment.

Key Points
  • Human-on-the-Bridge (HOB) places expert curation upstream — domain context, Red-Team Traps, Juror Personas — before testing begins.
  • ProofAgent Harness runs multi-turn adversarial evaluations with trace capture and multi-juror scoring across 23,500 agent turns.
  • HOB caught failures like phantom tool-calls, policy drift, manipulation paths, and safe but non-resolving refusals that benchmarks miss.

Why It Matters

Enables scalable, human-curated agent evaluation without large evaluator models, improving safety and reliability of deployed AI.

📬 Get the top 10 AI stories daily