Back to today's list

ZhuLong: Execution-Grounded LLM Agent for EDA Scripting with Offline API Self-Exploration

Yang Liu, Shiwei Hou, Xiyuan Chen, Yu Wang, Sen Yuan, Qirui Gan, Shao You, Feifan Chen, Wencheng Li, Shuyang Hu, Yongzhou Liu, Emma Xia, Xiaojing Lu, Hao Wang, Fan Xu, Yanfeng Li

Published Aug 11, 2026Featured #2In the daily list Aug 12, 2026
Daily score77.6
Editorial review8.2
Relevance0.458
Freshness0.722

Why It Matters

What makes this one worth your time

This work addresses a critical bottleneck in EDA scripting, providing a practical solution that enhances the usability of LLMs in complex environments, which is crucial for engineers in the field.

ZhuLong significantly improves EDA scripting by integrating execution-grounded learning with API self-exploration.

Summary

The paper introduces ZhuLong, an execution-grounded LLM agent designed for EDA scripting that utilizes API retrieval, documentation inspection, and sandbox execution, along with an offline API self-exploration mechanism to enhance performance on real-world tasks.

Key contributions

  • Development of ZhuLong, an execution-grounded LLM coding agent for EDA scripting.
  • Introduction of an offline API self-exploration mechanism for inferring undocumented API behaviors.
  • Evaluation on a benchmark of 158 real-world tasks, demonstrating significant performance improvements over existing LLMs.

Notable insights

  • The integration of sandbox execution as a performance driver highlights the importance of real-time feedback in LLM applications.
  • The offline API self-exploration mechanism demonstrates a novel approach to inferring undocumented API behaviors, which could be applicable in other domains.

Possible limitations

  • Not stated in the abstract.

Abstract

arXiv:2608.07925v1 Announce Type: new Abstract: EDA scripting with tool-specific, often undocumented APIs remains a long-tail bottleneck that existing LLMs fail to address. This paper presents ZhuLong, an execution-grounded LLM coding agent for PyAether and SKILL that combines API retrieval, documentation inspection, and sandbox execution via unified MCP tools, augmented by an offline API self-exploration mechanism that infers undocumented API behaviors through counterfactual experimentation. We evaluate ZhuLong on EDA-Eval-PyAether, a benchmark of 158 real-world tasks with assertion-based execution, where the complete system achieves 78.5% Pass@1 in the commercial Empyrean Aether environment, substantially outperforming a pure LLM baseline (23.6%). Ablation studies identify sandbox execution as the dominant performance driver (41.2 pp drop when removed), with the self-exploration mechanism contributing an additional 3.2 pp accuracy gain and a 22.1% reduction in per-task tool calls. On 20 interactive tasks involving unsaved layouts and schematics, ZhuLong achieves 60.0% Pass@1 for PyAether and 50.0% for SKILL.