SkillRevise: Improving LLM-Authored Agent Skills via Trace-Conditioned Skill Revision
Yuxuan Liu, Zhaochen Su, Lingyun Xie, Yuhao Zhang, Qing Zong, Jiahe Guo, Zhongwei Xie, Yiyan Ji, Yauwai Yim, Hongyu Luo, Xiyu Ren, Ruan Chenyu, Haoran Li, Yangqiu Song
Why It Matters
What makes this one worth your time
Improving the skill execution of LLM agents in cold-start scenarios can significantly enhance their practical utility and adaptability across different environments and tasks.
SkillRevise enhances LLM agent skills by iteratively refining initial skills based on execution evidence.
Summary
The paper introduces SkillRevise, a framework for refining initial skills of LLM agents by diagnosing defects from execution evidence and applying execution-anchored edits. It aims to improve skill performance in cold-start settings where only imperfect skills are available, and demonstrates significant improvements in success rates across benchmarks and LLMs.
Key contributions
- Proposes an execution-grounded framework for skill refinement in LLM agents.
- Demonstrates substantial improvement in skill success rates across multiple benchmarks.
- Shows that revised skills can transfer across different executors and task environments.
Notable insights
- SkillRevise uses execution evidence to diagnose and repair skill defects, which is a practical approach to improving skill performance.
- The framework retains the first verifier-passing skill within a revision budget, balancing performance improvement with resource constraints.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2606.01139v4 Announce Type: replace Abstract: Agent skills are procedural artifacts that enable LLM agents to execute workflows, verify constraints, and recover from failures. Existing self-evolving methods refine skills using accumulated trajectories. However, they struggle in cold-start settings, where only an initial, imperfect skill is available. Consequently, skill construction defaults to expert authoring or one-shot LLM generation. Expert-authored skills are costly and may not align with how LLM agents actually execute tasks, while one-shot generated skills can be syntactically well formed yet behaviorally weak. To bridge this gap, we propose SkillRevise, an execution-grounded framework designed to iteratively refine these initial skills. SkillRevise diagnoses skill defects from execution evidence, retrieves relevant repair principles from a general memory, and applies execution-anchored edits. By re-executing candidates and measuring empirical utility, it retains the best observed skill within the revision budget. Evaluated across three main benchmarks, two domain-specific studies, and six LLMs, SkillRevise substantially outperforms one-shot baselines, improving the base agent's success rate on SkillsBench from 36.05% to 61.63%. Furthermore, the revised skills transfer across both executors and task environments, suggesting that SkillRevise captures reusable procedural knowledge beyond any single executor. Our code is available at https://github.com/HKUST-KnowComp/skillrevise.