Back to today's list

Benchmarking LLM Competence on Logical Inference over Probability Operators

Nayera Hasan, Jack Greff, Alvin Grissom II

Published Aug 5, 2026Featured #7In the daily list Aug 5, 2026
Daily score60.7
Editorial review6.8
Relevance0.455
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding and improving logical inference in language models is crucial for their application in high-stakes domains like medicine and law.

A new benchmark reveals biases in language models' logical inference over probability operators.

Summary

The paper introduces a benchmark for evaluating large language models on logical inference tasks involving probability operators, using a dataset of 14,320 English prompts across various inference templates. It evaluates 29 models and finds that most exhibit biases and only a few exceed random chance in accuracy.

Key contributions

  • Introduction of a benchmark for logical inference over probability operators.
  • Evaluation of 29 language models on the benchmark, revealing systematic biases.

Notable insights

  • The study highlights systematic answer biases in language models, independent of logical form.
  • The benchmark includes variations in question form, verb phrases, and demographic attributes to test biases.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2607.27405v3 Announce Type: replace-cross Abstract: Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions of uncertainty are necessary for not only everyday conversations but also for high-stakes domains such as medicine and law. While large language models are increasingly evaluated on logical reasoning tasks, disentangling principled, symbolic reasoning from clever surface-level pattern matching is fraught with difficulty. We introduce a benchmark for reasoning over probability operators--inference over sentences with gradable epistemic modals (e.g., probably, might, must) containing 14,320 procedurally-generated English prompts across fifteen inference templates, systematically varying question form, negation strategy, and surface content. Evaluating 29 models, we find that most show answer biases independent of the logical form, a systematic preference for Yes or No. We summarize this with a competence floor: the worse of a model's accuracy on Yes-correct and No-correct items. Only 9 of 29 models exceed random chance. We also test variations in question form, verb phrases/activity, and both the gender and origin of names used in the prompts, finding biases across every axis.