Back to today's list

Tool-Making and Self-Evolving LLM Agents in Low-Latency Systems

Kalle Kujanp\"a\"a, Ning Liu, Shahnawaz Alam, Yeshwanth Reddy Sura, Tianyu Yang, Kristina Klinkner, Shervin Malmasi

Published Jul 10, 2026
Editorial review7.2
Relevance0.488
Freshness0.400

Why It Matters

What makes this one worth your time

This approach can make LLM systems more efficient and reliable, which is crucial for real-time applications in industrial settings.

Tool-making pipeline for LLM agents reduces latency and error rates in production systems.

Summary

The paper proposes a tool-making pipeline for LLM agents that compiles repeated procedural steps into validated tools, reducing latency and error rates in production systems. This approach was tested in a Fulfillment Center alarm-triage system, showing significant improvements in latency and error reduction.

Key contributions

  • Development of a tool-making pipeline for LLM agents.
  • Demonstrated latency and error rate improvements in a real-world alarm-triage system.
  • Introduction of versioned tools for better auditability and system analysis.

Notable insights

  • Compiling repeated procedural steps into tools reduces run-to-run variance and improves system reliability.
  • Versioned tools enhance auditability and help identify specification gaps and data drift.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2607.08010v1 Announce Type: new Abstract: Production LLM agents often waste latency and reliability by regenerating code for the same procedural steps on every request. We replace this inference-time coding loop with an agentic tool-making pipeline that compiles repeated SOP steps into validated, versioned tools before deployment. The tool-maker grounds synthesis in the live environment as it collects execution traces, observes backend schemas and values, generates candidate tools, and repairs them against labeled cases. At runtime, the production agent calls these tools directly and falls back to code generation only when needed. We deploy the approach in a Fulfillment Center alarm-triage system, where an agent diagnoses alarms against a 44-node SOP over heterogeneous metric backends. In production, tool calls reduce p50 latency by 42%. On 1,500 historical alarms, they reduce end-to-end error rate by up to 53% by suppressing run-to-run variance in repeated steps. Because tools return compact structured verdicts, they also enable a simpler direct-call architecture, reducing p50 latency by a further 62% in a controlled ablation. Versioned tools also improve auditability and expose specification gaps and upstream data drift. Our results show that self-evolving agents can make industrial LLM systems faster, more reliable, and easier to operate.