Research & Papers

TRACE Bench evaluates roleplay AI with 99.91% checklist coverage

Roleplay AI testing jumps from 73.74% to 99.91% coverage with traceable checklists.

Deep Dive

Roleplay evaluation has traditionally relied on single holistic scores that obscure which role requirements were met or missed. TRACE Bench, introduced by Jiahui Zhang and nine colleagues, replaces black-box scoring with a task-driven agentic checklist framework. Each role profile is decomposed offline into a fixed checklist. During evaluation, a User Agent engages the target roleplay model in natural conversation while privately updating checklist states based on the model's responses. This design allows every score to be traced back to specific checklist items and supporting dialogue turns, providing transparent, evidence-based judgments rather than vague impressions.

In cross-validation against MiniMax's Role-play Benchmark M2 free-dialogue transcripts, TRACE Bench covered only 73.74% of key role-profile points with released free-chat data, but TRACE Bench achieved 99.91% coverage in fewer turns. The framework proved robust across repeated runs and User Agent replacements, showing stable rankings. Across 26 models, it reports overall rankings, capability breakdowns, and checklist traces. It also supports Closed-Loop Benchmark Evolution, distilling verification methods proven effective in failed traces so future evaluations can more reliably elicit and examine failure modes. This makes TRACE Bench a practical tool for auditing and improving roleplay AI systems.

Key Points
  • TRACE Bench decomposes role profiles into fixed checklists, enabling traceable, evidence-based scoring instead of black-box holistic scores.
  • It boosts key role-profile coverage to 99.91% versus 73.74% for MiniMax's M2 free-dialogue transcripts, in fewer turns.
  • Evaluated across 26 models with stable rankings, and supports Closed-Loop Benchmark Evolution for eliciting failure modes.

Why It Matters

Roleplay AI evaluation moves from black-box scores to auditable, checklist-driven traces, improving reliability and diagnosis.

📬 Get the top 10 AI stories daily