Back to today's list

Human Grounded Evaluation of Large Language Models for Optical Network Automation

Kiarash Rezaei, Omran Ayoub, Paolo Monti, Carlos Natalino

Published Jul 22, 2026
Editorial review6.8
Relevance0.512
Freshness0.000

Why It Matters

What makes this one worth your time

This work is relevant for AI engineers and researchers interested in optimizing the use of LLMs in network automation, providing a method to balance explanation quality with computational efficiency.

HuGLEN offers a scalable method to evaluate and rank LLMs for optical network automation.

Summary

The paper introduces HuGLEN, an evaluation pipeline that combines LLM-as-a-judge with expert ratings to assess and rank large language models for optical network automation tasks, specifically focusing on translating XAI model outputs for QoT estimation into operator-friendly explanations.

Key contributions

  • Development of the HuGLEN evaluation pipeline.
  • Introduction of a quality efficiency score (QES) for ranking LLMs.
  • Application of HuGLEN to the optical network QoT estimation task.

Notable insights

  • Using LLM-as-a-judge can reduce the human-labeling burden in model evaluation.
  • A medium-sized LLM can achieve a better trade-off between quality and efficiency than larger models.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2607.18068v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly adopted for network automation, yet their output quality and inference cost can vary substantially across LLM families. We present HuGLEN, a stepwise evaluation pipeline that uses an LLM-as-a-judge together with a small set of expert ratings to enable scalable and reproducible comparison of candidate LLMs, and to rank them using a quality efficiency score (QES). We demonstrate HuGLEN for translating outputs from an explainable artificial intelligence (XAI) model for the optical network quality of transmission (QoT) estimation task into operator-friendly explanations. Our results show that a medium-sized LLM (12B parameters) achieves the highest QES, indicating the best trade-off between explanation quality and efficiency. Overall, HuGLEN reduces the human-labeling burden while supporting consistent model selection for operator-facing automation tasks.