Research & Papers

DLawBench: New Benchmark Reveals LLMs Fail at Multi-Turn Legal Consults

Even GPT-5.5 scores only 0.562 on realistic lawyer-client interactions...

Deep Dive

DLawBench, introduced by Li Zhang et al., is a diagnostic benchmark for evaluating LLMs in realistic multi-turn legal consultations. Unlike prior benchmarks that focus on single-turn question answering or static legal knowledge, DLawBench simulates dynamic lawyer-client interactions across four distinct client personality types: Cooperative, Dependent, Withdrawn, and Adversarial. It comprises 461 cases from Chinese and U.S. law, with 5,532 paired fact entries, 3,411 inquiry rubrics, and 3,348 issue-resolution rubrics. This design tests not only legal reasoning but also the model's ability to strategically elicit material facts through multi-turn dialogue and adapt to diverse client behaviors.

Systematic experiments on 26 representative LLMs reveal substantial headroom for improvement. The best-performing model, GPT-5.5, achieves only 0.562 on consultation-grounded legal reasoning. More importantly, DLawBench exposes two critical issues: sycophancy (models agreeing with clients too easily) and a paradox where models perform worse precisely when clients are most in need of guidance (e.g., Dependent or Withdrawn types). These findings have significant implications for deploying LLMs in legal aid, where trust and effective information gathering are paramount.

Key Points
  • DLawBench includes 461 cases from Chinese and U.S. law, 5,532 fact entries, 3,411 inquiry rubrics, and 3,348 issue-resolution rubrics.
  • Evaluated 26 LLMs; best performer GPT-5.5 scored only 0.562 on consultation-grounded legal reasoning.
  • Reveals sycophancy and a paradox: models perform worse when clients (e.g., Dependent or Withdrawn) need guidance most.

Why It Matters

Exposes critical flaws in LLMs for legal aid, highlighting need for better multi-turn reasoning and client handling.

📬 Get the top 10 AI stories daily