Research & Papers

RELIANCE dataset reveals 60% accuracy in TikTok reproductive health info, exposes LLM fact-checking gaps

New expert-annotated dataset from TikTok reveals critical gaps in LLM fact-checking for reproductive health

Deep Dive

Social media platforms like TikTok have become a primary source of health information, yet studies highlight widespread inaccuracies. As LLM providers (e.g., Grok on X, Perplexity on WhatsApp) integrate AI fact-checking, rigorous evaluation is essential—especially in sensitive areas like reproductive health. Researchers from multiple institutions present RELIANCE, a dataset curated from 336 TikTok videos covering 56 clinician-reviewed pregnancy and postpartum queries. Three expert clinicians in Obstetrics, Gynecology, and Internal Medicine annotated 409 sentences, creating a benchmark for both analyzing the health information landscape and evaluating LLM fact-checking capabilities.

The dataset reveals that nearly 60% of sampled health information is accurate, but LLM evaluations expose a critical 15% gap between assessing individual claims versus the entire video's content. This suggests current AI fact-checking may miss nuanced misinformation. Accepted at the ACM KDD 2026 Datasets and Benchmarks Track, RELIANCE provides open-source code and data to help the machine learning community develop safer LLM applications for health, extend analysis to other platforms and languages, and assist health professionals in understanding social media's information ecosystem.

Key Points
  • Dataset includes 409 clinician-annotated sentences from 336 TikTok videos across 56 pregnancy/postpartum queries.
  • Nearly 60% of health information in sampled videos was accurate, per expert review.
  • LLM fact-checking shows a 15% performance gap between claim-level and full-content evaluation.

Why It Matters

Essential dataset for improving LLM fact-checking in reproductive health—a critical area where errors cause serious harm.

📬 Get the top 10 AI stories daily