Research & Papers

Synthetic Rationale Data Hurts Alzheimer's Prediction in LLMs

Teaching AI to explain its reasoning actually makes it worse at diagnosing diseases.

Deep Dive

A new paper from researchers (Buxin Su et al.) tests a widely held assumption in clinical AI: that training language models to generate rationales (explanations) for their predictions improves performance. Using five-year Alzheimer's disease and related dementias (ADRD) prediction from longitudinal health records, the team ran a massive controlled experiment with 504 configurations across model families and data scales. The result is clear and counterintuitive: supervised fine-tuning (SFT) with synthetic rationale data consistently and substantially hurts prediction accuracy compared to standard label-only fine-tuning. The degradation persists even with reasoning-oriented base models and across different amounts of training data.

Crucially, the problem isn't poor rationale quality. Human expert annotation confirms the generated rationales are medically accurate and faithfully grounded in patient-specific evidence. In fact, using the same rationales as inference-time demonstrations (few-shot examples) actually improves performance. The researchers identify the root cause as a structural conflict between narrative plausibility (telling a coherent story) and discriminative optimization (making accurate predictions). This work challenges the prevailing approach of adding explanation training to clinical LLMs and calls for more precise understanding of when rationale-based supervision helps versus harms.

Key Points
  • Rationale-based SFT degrades Alzheimer's prediction across 504 configurations, consistently worse than label-only fine-tuning.
  • Human experts confirm the synthetic rationales are medically accurate, but using them as training targets still hurts performance.
  • Root cause is a structural conflict between narrative coherence and discriminative accuracy in clinical prediction tasks.

Why It Matters

Challenges the assumption that adding explanation training improves clinical LLMs, guiding more responsible AI development in healthcare.

📬 Get the top 10 AI stories daily