Research & Papers

QIAS 2026: LLMs fail Islamic inheritance reasoning with 12,500 cases

Only 16 teams could tackle the MAWARITH benchmark's legal and numerical challenges.

Deep Dive

Researchers from multiple institutions organized QIAS 2026, a shared task designed to evaluate large language models on complex reasoning in Islamic inheritance law. The task required full end-to-end processing: from parsing natural language case descriptions to identifying eligible heirs and calculating their exact shares according to Sharia rules. The evaluation used the MAWARITH dataset, a collection of 12,500 Arabic inheritance cases annotated with intermediate reasoning steps and final answers. Systems were scored using MIR-E, a multi-step metric that measures performance across each stage of inheritance reasoning.

A total of 16 teams participated, exploring prompting-based methods, retrieval-augmented generation, and fine-tuning strategies. Despite these diverse approaches, results indicate that Islamic inheritance remains a highly challenging domain for current LLMs. The main difficulties lie in precise legal interpretation (e.g., identifying which relatives qualify as heirs) and structured numerical reasoning (e.g., correctly applying fractional share calculations). The task highlights the gap between general language understanding and domain-specific, rule-bound reasoning—critical for applications in legal and religious contexts.

Key Points
  • QIAS 2026 used the MAWARITH benchmark with 12,500 annotated Arabic inheritance cases for end-to-end reasoning.
  • Sixteen teams competed using prompting, RAG, and fine-tuning; none achieved high scores on legal interpretation stages.
  • The MIR-E evaluation metric tracks performance across multiple reasoning steps, revealing LLMs' weaknesses in structured numerical computation.

Why It Matters

This benchmark exposes critical LLM limitations in rule-based legal reasoning, affecting adoption in high-stakes domains like religious law and finance.

📬 Get the top 10 AI stories daily