Research & Papers

GIScholarBench reveals LLMs overconfident in academic tasks

Claude, Gemini, and ChatGPT confidently produce wrong DOIs and citations

Deep Dive

A new benchmark, GIScholarBench, reveals that large language models exhibit systematic overconfidence when used for academic research tasks. Built from 10,865 papers across 25 core GIScience journals published between 2020 and 2025, the benchmark tests LLMs on three increasingly complex tasks: metadata retrieval, literature linking, and research direction generation. The researchers evaluated three leading models—Claude Sonnet 4.5, Gemini 3, and ChatGPT 5.3—through their native web interfaces under real-world conditions.

Results show task-invariant overconfidence across all models. In metadata retrieval, ChatGPT 5.3 achieved the highest accuracy, but all models confidently generated definitive-sounding titles and DOIs even when predictions were wrong. In literature linking, Claude Sonnet 4.5 recovered the most references, yet exhibited a clear gap between top-ranked results and longer citation lists, indicating unreliable citation expansion. For research direction generation, AI outputs had lower topic coverage, higher rates of novel misses, and lower semantic diversity compared to actual future papers. The authors conclude that overconfidence manifests differently per task: factual overgeneration in retrieval, unreliable citation expansion in linking, and overconfidence in output completeness during ideation.

Key Points
  • GIScholarBench includes 10,865 papers from 25 core GIScience journals (2020-2025).
  • ChatGPT 5.3 had highest metadata accuracy but still fabricated DOIs; Claude Sonnet 4.5 retrieved most references but expanded citations unreliably.
  • AI-generated research directions showed 25% lower topic coverage and higher novelty miss rates vs. real future papers.

Why It Matters

Researchers relying on LLM outputs risk propagating confidently wrong facts, citations, and research directions.

📬 Get the top 10 AI stories daily