Research & Papers

CIKM 2026 study: Persona conditioning skews LLM relevance assessors in IR evaluation

Five assessor personas on six LLMs shift judgment strictness, not relevance—threatening system rankings

Deep Dive

LLMs are increasingly replacing human annotators in information retrieval (IR) evaluation, but new research shows their judgments are surprisingly sensitive to the persona they're given. In a paper accepted at CIKM 2026, Samaneh Mohtadi, Pietro Bernardelle, Joel Mackenzie, and Gianluca Demartini introduced "persona conditioning" as a diagnostic probe—essentially instructing an LLM to adopt a specific assessor role (e.g., intent interpreter, domain expert, evidence verifier) before scoring relevance. Running five persona types against a standard UMBRELA baseline across six LLM backbones and two benchmarks (TREC DL20 and RAG24), they discovered the sensitivity is structured, not chaotic: persona shifts rarely flip relevance scores wholesale, but they systematically alter judgment strictness, evidential thresholds, and interpretation emphasis.

The more critical finding is at the system level. High-capacity LLMs preserved overall system-ranking agreement regardless of persona, while smaller models amplified persona-induced instability, meaning evaluation results from weaker assessors can be materially skewed by phrasing. Local rank-displacement analysis showed disturbances concentrate on specific retrieval system types—neural ranking/reranking systems in DL20 and RAG-oriented pipelines in RAG24. Notably, the persona's source (PersonaHub vs. NVIDIA Nemotron-Personas-USA) mattered less than the role itself and model capacity. The authors position persona-conditioned judging as a controlled sensitivity probe: by varying assessor framing, researchers can identify which IR systems produce evaluation outcomes that are dangerously fragile to prompt wording—without needing new benchmarks or human re-annotation.

Key Points
  • Tested 5 assessor personas (intent, expertise, contrastive, evidence, global quality) vs UMBRELA baseline across 6 LLM backbones on TREC DL20 and RAG24
  • High-capacity models preserve system-ranking agreement; smaller models amplify persona-induced instability
  • Sensitivity localizes to neural ranking/reranking on DL20 and RAG pipelines on RAG24, with persona source mattering less than role and model size

Why It Matters

Persona conditioning gives IR teams a cheap, reproducible way to stress-test whether LLM-based evaluation rankings hold up under prompt variation.

📬 Get the top 10 AI stories daily