New study reveals LLM self-consistency paradox: more consistent, more mistakes
Testing 10 frontier models across 491 concepts reveals a dangerous trade-off in agentic pipelines.
A new measure, 'generator-evaluator self-consistency,' is proposed to test LLMs that self-evaluate without external verification. Studying 10 frontier models across 491 concepts, the researchers find that higher self-consistency correlates with greater vulnerability to mistakes in a clinical setting with physician-validated errors (Proniakin et al., 2025). This 'consistency dilemma' shows even consistent models may be unsafe to deploy in agentic pipelines.
- New metric 'generator-evaluator self-consistency' tests if LLMs apply concepts the same way when generating and evaluating outputs across 10 frontier models and 491 concepts.
- Higher self-consistency correlates with greater vulnerability to mistakes in a clinical setting using physician-validated error data (Proniakin et al., 2025).
- The 'consistency dilemma': operationally useful self-consistency can make models unsafe for agentic pipelines because they repeat the same systematic errors.
Why It Matters
This challenges the safety of self-evaluating AI agents — consistent models aren't necessarily trustworthy, especially in high-stakes domains like healthcare.