Back to today's list

WuYuEval: A Multi-Level Benchmark for Large Language Models in Solid Waste Management

Yi Zhang, Hongyang Wang, Zheng Hao Leong, Zihao Wu, Kaijun Lin, Zhixing Pan, Qixun Huangfu, Wei Ren, Wenyan Wu, Fangyun Wang, Wenting Yu, Hengyu Lin, Muling Yang, Zongguo Wen

Published Aug 11, 2026
Editorial review6.8
Relevance0.478
Freshness0.000

Why It Matters

What makes this one worth your time

This benchmark provides a structured way to evaluate and improve LLMs' capabilities in the specialized field of solid waste management, which is crucial for developing models that can handle domain-specific challenges.

WuYuEval is a benchmark for assessing LLMs in solid waste management tasks.

Summary

The paper introduces WuYuEval, a benchmark designed to evaluate large language models (LLMs) in the domain of solid waste management (SWM). It includes a Foundation Module with multiple-choice questions and an Expert Module with scenario-based questions. The benchmark assesses LLMs' performance in foundational knowledge, domain reasoning, and expert decision-making, revealing varying performance across models and tasks.

Key contributions

  • Introduction of WuYuEval, a multi-level benchmark for LLMs in SWM.
  • Development of a Foundation Module with 4,590 multiple-choice questions and an Expert Module with 247 open-ended questions.
  • Use of a novel scoring method combining LLM-as-a-Judge and Elo-based comparison.

Notable insights

  • The use of anchor-calibrated LLM-as-a-Judge scoring combined with Elo-based pairwise comparison for expert tasks.
  • Visible deliberation in reasoning improves performance only when anchored to units, assumptions, and constraints.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2608.07529v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as technical assistants, but their competence in solid waste management (SWM) remains difficult to assess because existing benchmarks emphasize general knowledge rather than professional decisions under engineering, environmental, and policy constraints. We introduce WuYuEval, a multi-level benchmark for evaluating LLMs in SWM across foundational knowledge, domain reasoning, and expert decision-making. After quality auditing, WuYuEval contains a Foundation Module with 4,590 closed-ended multiple-choice questions across six task types and eight domain categories, together with an Expert Module with 247 scenario-based open-ended questions involving multi-objective optimization, constraint trade-offs, and system design. For expert tasks, we combine anchor-calibrated LLM-as-a-Judge scoring with Elo-based pairwise comparison. Across 33 LLMs, performance varied widely. The leading model reached 94.64\% accuracy on the Foundation Module, but average accuracy still fell from 84.14\% on easy questions to 42.50\% on hard questions, with lower performance concentrated in calculation, experimental design, urban planning, and open-ended expert tasks. Reasoning-oriented Thinking modes improve most matched model pairs after auditing, but the gains depend on baseline capability and are not uniformly positive. These results suggest that visible deliberation helps only when it remains anchored to units, assumptions, and engineering constraints; otherwise, it may drift from decisive answer boundaries. WuYuEval therefore provides both an evaluation resource and an empirical basis for developing SWM-oriented foundation models with professional reasoning chains and explicit constraint control.