Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary
S. Ashwin Hebbar, Peiyao Sheng, Sewoong Oh, Pramod Viswanath
Why It Matters
What makes this one worth your time
This work is significant for improving the reliability of AI-generated content in chess, which can enhance educational tools for both experts and novices.
ACT-Eval enhances the evaluation of LLM-generated chess commentary by addressing factual inaccuracies.
Summary
The paper introduces ACT-Eval, an evaluation framework for assessing the factual correctness and conceptual coverage of chess commentary generated by large language models, using expert-annotated references and engine-supported tools.
Key contributions
- Development of ACT-Eval, a novel evaluation framework for chess commentary.
- Creation of a benchmark dataset of 325 position-move pairs with expert-verified gold atoms.
- Introduction of a five-class error taxonomy for better categorization of factual inaccuracies.
Notable insights
- The framework's decomposition of chess commentary into atomic claims allows for more granular evaluation of LLM outputs.
- The correlation of ACT-Eval's coverage scores with human assessments suggests a robust method for evaluating strategic completeness.
Possible limitations
- The abstract does not address potential biases in expert annotations or the representativeness of the benchmark dataset.
- Limited coverage of expert strategic and tactical ideas across all models suggests room for improvement.
Abstract
arXiv:2608.04240v1 Announce Type: cross Abstract: Superhuman game engines in domains like chess have made expert-level evaluations easily accessible, yet they communicate what is true without the natural-language explanations that make such expertise educationally useful to experts and non-experts alike. Large language models could, in principle, bridge this gap, but they frequently hallucinate due to limited domain-specific knowledge, and standard reference-based or LLM-as-a-judge frameworks cannot reliably detect these errors. In this work, we present ACT-Eval, an evaluation framework that decomposes chess commentary into atomic claims and routes them to engine-supported tools and expert-annotated gold references to assess factual correctness, conceptual coverage, and move-quality judgment. We release a benchmark of 325 position--move pairs spanning pedagogical, tournament, and critical positions, including 125 positions with expert-verified gold atoms and a five-class error taxonomy. Evaluating leading proprietary and open-weight models, we find that factual hallucinations remain pervasive in chess commentary: GPT-5.4 without tools produces incorrect sub-claims 22.0% of the time, while smaller open-weight models exceed 40%. Although tool augmentation substantially improves factual correctness and move-quality assessment, coverage of expert strategic and tactical ideas remains limited across all models. Human calibration shows that ACT-Eval's factual judgments fall within the observed range of inter-human agreement, while its coverage scores correlate strongly with human assessments of strategic completeness.