CombEval: A Framework for Evaluating Combinatorial Counting in Large Language Models
Yuxu Zhou, Ond\v{r}ej Ku\v{z}elka, Yuyi Wang, Yuanhong Wang, Yi Chang
Why It Matters
What makes this one worth your time
Understanding the limitations of LLMs in combinatorial reasoning can guide improvements in their design and application, particularly in tasks requiring precise counting and constraint handling.
CombEval is a new benchmark for testing LLMs on combinatorial counting tasks.
Summary
The paper introduces CombEval, a dynamic benchmark designed to evaluate combinatorial counting capabilities in large language models (LLMs). It uses a typed Cofola specification to generate natural-language counting problems with solver-verified answers, allowing for systematic variation in problem parameters. The study evaluates 11 LLMs and identifies their brittleness in handling ordered objects, indistinguishable elements, and complex constraints.
Key contributions
- Development of CombEval, a dynamic benchmark for combinatorial counting.
- Evaluation of 11 LLMs, highlighting their limitations in combinatorial reasoning.
- Provision of a publicly available codebase and benchmark suite.
Notable insights
- The use of a typed Cofola specification allows for controlled generation of counting problems with exact solutions.
- Error analysis reveals specific weaknesses in LLMs' handling of ordered objects and constraint interpretation.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2606.19788v2 Announce Type: replace Abstract: We present CombEval, a dynamic benchmark for evaluating combinatorial counting in large language models. CombEval represents each problem as a typed Cofola specification over entities, combinatorial objects, object dependencies, and constraints, enabling controlled generation of natural-language counting problems with exact solver-verified answers. Unlike static collections, CombEval supports systematic variation of object type, entity scale, constraint count, and reasoning depth. We evaluate 11 LLMs under direct and code-augmented settings and find that models remain brittle on ordered objects, indistinguishable elements, relatively positional constraints, and nested object dependencies. Error analysis further identifies failures in constraint interpretation and counting principles. CombEval provides a diagnostic testbed for studying when and why LLMs fail at combinatorial reasoning. The code and generated benchmark suites are publicly available at https://github.com/YuxuZhou-CN/combination-problem-generation.