Back to today's list

Benchmarking LLM Competence on Logical Inference over Probability Operators

Nayera Hasan, Jack Greff, Alvin Grissom II

Published Jul 31, 2026Featured #3In the daily list Aug 1, 2026
Daily score72.7
Editorial review7.5
Relevance0.458
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding LLMs' reasoning capabilities is crucial for applications in high-stakes fields like medicine and law, where accurate inference is essential.

A new benchmark reveals biases in LLMs' logical reasoning with probability operators.

Summary

The paper introduces a benchmark for evaluating large language models on logical inference involving probability operators, presenting a dataset of 14,320 prompts and analyzing biases in model responses.

Key contributions

  • Introduction of a large-scale benchmark for logical inference over probability operators.
  • Evaluation of 29 models, revealing systematic biases in responses.
  • Analysis of the impact of question form and demographic factors on model performance.

Notable insights

  • The benchmark systematically varies question form and negation strategy, providing a comprehensive evaluation framework.
  • The identification of a 'competence floor' highlights the need for improved model training in logical reasoning.

Possible limitations

  • Not stated in the abstract.

Abstract

arXiv:2607.27405v1 Announce Type: new Abstract: Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions of uncertainty are necessary for not only everyday conversations but also for high-stakes domains such as medicine and law. While large language models are increasingly evaluated on logical reasoning tasks, disentangling principled, symbolic reasoning from clever surface-level pattern matching is fraught with difficulty. We introduce a benchmark for reasoning over probability operators--inference over sentences with gradable epistemic modals (e.g., probably, might, must) containing 14,320 procedurally-generated English prompts across fifteen inference templates, systematically varying question form, negation strategy, and surface content. Evaluating 29 models, we find that most show answer biases independent of the logical form, a systematic preference for Yes or No. We summarize this with a competence floor: the worse of a model's accuracy on Yes-correct and No-correct items. Only 9 of 29 models exceed random chance. We also test variations in question form, verb phrases/activity, and both the gender and origin of names used in the prompts, finding biases across every axis.