Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents
Tao Wang, Qihao Yang, Rongjiao Liang, Lianghong Lin, Haitao Wang, Xinyu Cao, Tianyong Hao
Why It Matters
What makes this one worth your time
This research is relevant for AI engineers and researchers interested in improving LLM capabilities for complex document reviews, potentially reducing reliance on costly human experts.
The paper develops a benchmark and framework to enhance LLMs for reviewing rule-intensive national standard documents.
Summary
The paper introduces GB/T-Bench, a benchmark for evaluating large language models (LLMs) in the rule-intensive review of national standard documents, specifically focusing on China's GB/T standards. It presents a hierarchical taxonomy for document review and a multi-agent framework called GB/T-Reviewer to improve LLM performance in this domain.
Key contributions
- Introduction of GB/T-Bench, a benchmark for structured review of national standard documents.
- Development of a hierarchical taxonomy covering various document review dimensions.
- Proposal of GB/T-Reviewer, a multi-agent framework to enhance LLM performance in rule-intensive document review.
Notable insights
- The use of a hierarchical taxonomy and a multi-agent framework to coordinate specialized skills for document review is a clever approach to improving LLM performance.
- The combination of deterministic rules and constrained LLM rewriting for generating counterexamples is an innovative method for creating a robust evaluation dataset.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2608.06312v1 Announce Type: new Abstract: Large language models (LLMs) increasingly support complex professional tasks, yet their capabilities in rule-intensive document review remain insufficiently evaluated. National standard documents, such as China GB/T standards, offer a representative testbed: they are lengthy, highly structured, and governed by explicit rules for scope, terminology, normative wording, and cross-section consistency. Existing benchmarks focus on domain knowledge and question answering, largely overlooking intrinsic quality review for professional documents. Such reviews rely heavily on human experts, making them costly and difficult to scale. To bridge this gap, we introduce GB/T-Bench, the first benchmark for the structured review of national standard documents. Its GB/T Review Taxonomy is a hierarchical schema covering document structure, scope alignment, normative modality, terminology consistency, and normative references, with 25 diagnosable error types. A controllable counterexample generation mechanism combines deterministic rules and constrained LLM rewriting to process 488 documents into 7,306 traceable review error instances for evaluation. We also develop a diagnosis-oriented evaluation protocol requiring exact matches on error location, review dimension, and error type, plus document-level coverage metrics. We further propose GB/T-Reviewer, a multi-agent framework that converts review knowledge into specialized skills and coordinates global inspection, targeted diagnosis, rule scanning, and result verification. Experiments with 14 mainstream LLMs reveal a substantial human-LLM gap: the strongest model achieves only 0.3280 CMCS versus 0.6640 for experts. GB/T-Reviewer raises the best CMCS to 0.5094, showing the value of structured skill coordination for rule-intensive document review. This work paves the way for trustworthy AI in standardization and other high-stakes document domains.