Anthropic-backed Conceptual Reasoning Index tests AI on unverifiable safety questions
Three new benchmarks challenge AI where empirical feedback is impossible—can models reason philosophically?
The Conceptual Reasoning Index (CRI), developed in collaboration with Anthropic, is a new benchmark suite designed to measure AI's conceptual reasoning capabilities—the kind of argumentation used in philosophy, AI futurism, and high-level AI safety work. The project compiles three benchmarks: LMCA, ACCoRD, and DTBench, which focus on reasoning about AI governance, alignment, and avoiding catastrophic cooperation failures. Unlike conventional benchmarks that rely on empirical or mathematical verification, these tasks deliberately lack a clear ground truth, forcing models to rely on structured argumentation and judgment. The team argues that AI risk mitigation often requires getting things right the first time, with no opportunity for trial-and-error learning.
CRI is available at conceptualreasoning.ai, where the team will publish results as new models and benchmarks are released. Researchers can request access to the primary conceptual dataset, LMCA, through an online form. The motivation is practical: current AI training depends on abundant data and reliable feedback, but many high-stakes safety decisions—such as which research agendas to prioritize or which values AI should hold—offer no empirical feedback loop. By quantifying how well models handle these unverifiable, long-horizon questions, CRI aims to identify gaps and potentially guide efforts to selectively improve AI's skills in reducing existential risk. The creators stress that this is an early step, with detailed blog posts arguing the case forthcoming.
- CRI aggregates three conceptual reasoning benchmarks: LMCA, ACCoRD, and DTBench, designed to evaluate AI on unverifiable, argumentation-heavy tasks.
- The project was developed in collaboration with Anthropic and is hosted at conceptualreasoning.ai, with LMCA dataset access available via a request form.
- Focuses on reasoning where no empirical or mathematical verification exists, such as AI governance, alignment, and long-horizon risk planning—critical for AI risk mitigation.
Why It Matters
If AI must help reduce existential risk, it needs to reason on unverifiable questions—this benchmark measures that gap.