LiveEvalBench: Toward Open-World Evaluation for Web Generation
Yiyao Wang, Zhen Wen, Yinghao Tang, Yixiao Fu, Lin Yuan, Xiaolau Zhang, Jun Zhou, Wei Chen
Why It Matters
What makes this one worth your time
This framework addresses the limitations of static evaluation methods for web generation, providing a more comprehensive and adaptable approach that aligns with real-world needs.
LiveEvalBench offers a dynamic framework for evaluating web generation with collaborative roles.
Summary
The paper introduces LiveEvalBench, a framework for evaluating web generation by treating it as an interactive and adaptive process rather than a static one. It involves a collaborative review workflow with roles like Build Engineer, Code Engineer, and UI Tester to assess frontend projects throughout their lifecycle. The framework allows for diverse implementations and incremental integration of new evaluation roles.
Key contributions
- Introduction of LiveEvalBench, an adaptive framework for web generation evaluation.
- Collaborative review workflow involving multiple evaluator roles.
- Adaptive protocol for handling diverse implementations.
Notable insights
- The framework uses a collaborative review workflow to evaluate web generation projects.
- It supports diverse implementations and incremental integration of new evaluation roles.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2608.03689v1 Announce Type: new Abstract: Large language models are increasingly capable of synthesizing executable frontend projects, yet existing benchmarks still treat web generation as a static evaluation problem. We argue that frontend artifacts demand a different paradigm: they are interactive rather than static, admit diverse yet equally valid implementations, and evolve faster than rigid pipelines can accommodate. To address these gaps, we present LiveEvalBench, an automated framework that reformulates web-generation evaluation as an agentic, adaptive, and extensible process. LiveEvalBench instantiates evaluation as a collaborative review workflow, in which a Build Engineer, a Code Engineer, and a UI Tester collectively gather evidence across the full lifecycle of a frontend project, from deployment and code inspection to browser-based interaction. To handle implementation diversity, an adaptive protocol couples shared rubrics for cross-model comparability with implementation-grounded criteria tailored to each artifact. The framework further supports incremental integration of new evaluator roles and assessment dimensions without pipeline redesign. Experiments across diverse real-world web-generation scenarios show that LiveEvalBench aligns closely with human expert judgment and provides fine-grained insights into frontier models' web generation capabilities. Code is available at https://github.com/wyysteelhead/LiveEvalBench