Research & Papers

New benchmark: 29 LLMs tested, only 9 beat chance on probability reasoning

LLMs fail at logic with 'probably' and 'might'—most just answer Yes

Deep Dive

A new paper from researchers at (affiliation not stated) introduces a benchmark to test whether large language models can perform logical inference over probability operators—words like 'probably,' 'might,' and 'must.' These expressions are common in everyday language and critical in high-stakes fields like medicine and law, where valid reasoning about uncertainty is essential. The benchmark, titled 'Benchmarking LLM Competence on Logical Inference over Probability Operators,' contains 14,320 procedurally-generated English prompts across 15 inference templates, systematically varying question form, negation strategy, and surface content.

The results are sobering. When the authors evaluated 29 models, they found most exhibit answer biases independent of logical form—a systematic preference for either 'Yes' or 'No' regardless of whether that answer is correct. Using a 'competence floor' metric (the lower of a model's accuracy on Yes-correct and No-correct items), only 9 of 29 models exceeded random chance. The team also tested variations in question wording, verb phrases/activity, and both the gender and origin of names in prompts, finding measurable biases across every axis. This suggests that models are often using surface-level heuristics rather than the underlying logic of uncertainty.

The benchmark highlights a significant gap in LLM reasoning: while models can handle straightforward factual logic, gradable epistemic modals remain a weak spot. The findings imply that current LLMs are unreliable for tasks that require nuanced probabilistic reasoning, which has direct implications for AI assistants in medical diagnosis, legal analysis, and risk assessment. The authors plan to release the benchmark publicly, allowing other researchers to further investigate and improve model performance on this challenging class of inference.

Key Points
  • 14,320 procedurally-generated prompts across 15 inference templates testing probability operators like 'probably' and 'might'
  • Only 9 of 29 LLMs exceeded random chance on the competence floor metric; most showed systematic Yes/No biases
  • Biases persist across question form, verb phrases, and name gender/origin, exposing surface-level heuristics in reasoning

Why It Matters

Highlights that LLMs poorly handle uncertainty logic, risking errors in medical, legal, and risk-critical applications.

📬 Get the top 10 AI stories daily