Research & Papers

Know2Guess Benchmark Exposes When LLMs Should Guess vs Abstain

1,200-item benchmark reveals Qwen2.5-3B best at knowing when to say 'I don’t know'

Deep Dive

A team of researchers from multiple institutions has released Know2Guess, a novel benchmark designed to rigorously evaluate whether large language models can reliably answer questions they should know and abstain from those they shouldn't—without conflating performance with data contamination or prompt idiosyncrasy. The benchmark contains 1,200 items across five domains, each with explicit abstention expectations and contamination-risk metadata. It uses dual parsing (strict and normalized) and locked answer-or-abstain prompts to isolate genuine knowledge boundaries from generic refusal behavior.

Testing FLAN-T5, Qwen2.5-Instruct, and Llama-3-Instruct models revealed that no model fully masters the answer-or-abstain task. Qwen2.5-3B-Instruct achieved the best overall reliability, but calibration remains poor and benign-item refusal persists. FLAN-T5 baselines showed weak productive abstention. The benchmark exposes a selective, incomplete transition from answering to abstaining even in strong instruction-tuned models. Prompt and parser robustness analyses preserved rankings, confirming Know2Guess is a usable, reproducible protocol for auditing how LLMs handle knowledge boundaries, with the dataset publicly available.

Key Points
  • Know2Guess benchmark includes 1,200 items across 5 domains (e.g., science, history) with explicit abstention expectations and contamination-risk metadata
  • Qwen2.5-3B-Instruct scored highest overall reliability, but answer-expected zones remain difficult and calibration is poor across all tested models
  • FLAN-T5 baselines showed weak productive abstention, while stronger models (Llama-3) exhibited selective but incomplete answer-to-abstain transitions

Why It Matters

Professionals relying on LLMs need to trust when models admit ignorance—this benchmark provides the first reproducible tool to audit that behavior.

📬 Get the top 10 AI stories daily