New study exposes critical benchmark gap for legal AI under EU AI Act
LLMs pass legal tests but can't do doctrinal reasoning—EU law demands it.
A new paper by legal scholar Michèle Finck (arXiv, June 2026) exposes a critical blind spot in the automation of EU law: current benchmarks only evaluate LLMs on paralegal tasks—like summarization or document review—not on the doctrinal legal reasoning that forms the interpretive core of legal work. Finck argues that although LLMs now produce legal text of at least median quality, no existing test can determine whether a model truly reasons like a lawyer. This measurement gap is not merely academic. The EU AI Act classifies AI used in the judicial domain as high-risk and mandates 'appropriate accuracy' as a binding requirement. Without a benchmark tailored to doctrinal reasoning, that legal standard lacks operational meaning. Regulators cannot assess compliance, and developers have no target to aim for. The paper positions itself as a call to action for the AI and legal communities to build and adopt a rigorous evaluation suite for doctrinal legal reasoning—something that currently does not exist. Finck's work bridges computer science and jurisprudence, highlighting how regulatory frameworks demand technical solutions that have not yet been invented. The findings have immediate implications for legal AI startups, law firms deploying LLMs, and EU policymakers tasked with enforcing the AI Act.
- LLMs achieve median-quality legal text but lack benchmarks for doctrinal legal reasoning—the core of legal analysis.
- EU AI Act's 'appropriate accuracy' requirement for high-risk judicial AI cannot be enforced without a doctrinal-reasoning benchmark.
- Paper calls for a new evaluation framework that measures interpretive reasoning, not just paralegal tasks.
Why It Matters
Without a doctrinal-reasoning benchmark, EU legal AI compliance is unmeasurable—risking regulatory failure and unsafe deployment.