Research & Papers

LLMs can ace legal exams while citing wrong laws, study finds

Up to 42.4% of correct legal answers miss the gold authority

Deep Dive

Legal benchmarks typically grade only whether an LLM picks the right final answer, ignoring whether the model actually grounds that answer in the correct statute. In a new arXiv paper, Hsien-Jyh Liao demonstrates that this oversight masks a systematic gap: models can look legally fluent without being legally accurate. Using 238 Taiwan bar-examination questions, each with a verified governing provision, Liao ran four LLMs under ordinary reasoning prompts that did not ask for citations. The models spontaneously produced authority markers anyway, allowing an automated joint audit of answer correctness and authority grounding.

The results show a clear dissociation between the two dimensions. In criminal law, 24.0–42.4% of valid responses were answer-correct but cited the wrong authority or no gold authority, while 15.2–21.7% were answer-incorrect yet did cite the correct provision. A separate statutory-retrieval probe and a permissive citation-abstention intervention confirmed that answer and citation behaviors can move independently at the output level. Because this mismatch occurs without adversarial prompting, existing answer-only scoring treats naturally occurring gold-authority misses as complete benchmark successes. A preliminary extension to PRC civil law also observed citation-unrequested authority marking, suggesting the issue is cross-jurisdictional. Liao argues that statute-grounded legal benchmarks should adopt joint answer–authority evaluation, and notes the failure is automatically measurable since statutory authority is structurally extractable and externally verifiable.

Key Points
  • Audited 238 Taiwan bar-exam items across four LLMs without citation-requesting prompts
  • Criminal-law responses: 24.0–42.4% answer-correct but missed gold authority; 15.2–21.7% answer-incorrect but cited it
  • Proposes joint answer–authority scoring for legal benchmarks, with a PRC civil-law pilot showing similar issues

Why It Matters

Legal AI benchmarks may overstate real competence, hiding wrong-law citations in seemingly correct answers.

📬 Get the top 10 AI stories daily