Research & Papers

CANDI-QA benchmark reveals LLMs struggle with niche domain questions

New dataset tests LLMs on medical and financial reasoning — most models fail badly.

Deep Dive

A team of researchers led by Megha Chakraborty has introduced CANDI-QA, a novel benchmark designed to evaluate large language models (LLMs) on question answering in highly specialized domains such as medical diagnostics and financial advisory. Unlike traditional QA datasets that test general knowledge, CANDI-QA focuses on contextual alignment, user awareness, and domain-specific understanding. The dataset is split into two categories: Information Assistance Questions (direct factual queries requiring precise extraction) and Applied Inference Questions (multi-hop reasoning tasks that demand situational inference). The team evaluated over ten diverse LLMs, ranging from compact open-source models to state-of-the-art proprietary systems.

To establish a robust baseline, the authors developed MTSS-Net, a lightweight neuro-symbolic framework that combines neural retrieval with rule-based reasoning. Their findings reveal that current LLMs struggle significantly with contextual grounding in niche domains, often failing to generate accurate or actionable answers. MTSS-Net outperformed pure LLM approaches on these tasks, highlighting the need for enhanced symbolic integration and context-aware architectures. The CANDI-QA benchmark aims to drive research toward more trustworthy AI for high-stakes applications, where mistakes can have serious real-world consequences.

Key Points
  • CANDI-QA features 2 categories: factual extraction and multi-hop reasoning, both expert-curated for medical/financial domains.
  • Over 10 LLMs tested, from compact open-source to top proprietary models; all struggled with contextual alignment.
  • MTSS-Net, a neuro-symbolic baseline combining retrieval and rule-based reasoning, outperformed pure LLMs on niche QA.

Why It Matters

This benchmark exposes critical gaps in LLM reliability for high-stakes fields like medicine and finance, guiding safer AI deployment.

📬 Get the top 10 AI stories daily