Research & Papers

Google's Gemini 2.5 Flash tops AI inheritance reasoning benchmark

Commercial LLMs trounce open-source in complex Islamic legal calculations

Deep Dive

A new paper from team PSL (Paris Sciences & Lettres) reveals a stark performance gap between commercial and open-source large language models in legal inheritance reasoning. The study, submitted to the QIAS 2026 Shared Task, tested models on Arabic Islamic inheritance cases requiring legal interpretation, multi-step reasoning, and precise numerical computation.

Google's Gemini 2.5 Flash achieved the lowest Mean Relative Error (MRE) of 0.989, significantly outperforming all other models. Commercial models consistently identified eligible heirs, applied exclusion rules correctly, and maintained logical coherence across complex dependency chains. In contrast, open-source models frequently faltered on dependent legal decisions and fractional share adjustments, producing unstable outputs.

The findings underscore a critical reliability issue for deploying open-source LLMs in high-stakes legal or regulatory contexts. As AI moves into specialized domains requiring both legal expertise and mathematical precision, the commercial vs. open-source divide may widen further—especially in languages and legal systems underrepresented in training data.

Key Points
  • Gemini 2.5 Flash achieved best performance with MRE of 0.989 on inheritance reasoning
  • Commercial models outperformed open-source in heir identification, exclusion rules, and multi-step consistency
  • Open-source models showed significant instability in fractional share adjustments and dependent legal decisions

Why It Matters

Reveals critical reliability gap between commercial and open-source LLMs for specialized legal reasoning.

📬 Get the top 10 AI stories daily