Beyond Component Testing: Validating Agentic AI Systems
Fabio Orazio Mirto, Luca D'Agati, Giuseppe Tricomi, Stefano Silvestri, Francesco Longo, Antonio Puliafito, Giovanni Merlino
Why It Matters
What makes this one worth your time
Understanding and validating the complex behaviors of agentic AI systems is crucial for their safe and trustworthy deployment in real-world applications.
The paper proposes a taxonomy for validating agentic AI systems beyond component testing.
Summary
The paper surveys 257 papers to address the validation challenges of agentic AI systems, proposing a five-dimension taxonomy to map current approaches and identify gaps in behavioral, safety, temporal, regulatory, and multi-agent concerns. It highlights the maturity of behavioral evaluation and the underdevelopment of other dimensions, using case studies to illustrate these issues in safety-critical settings.
Key contributions
- Proposes a five-dimension taxonomy for agentic AI validation.
- Synthesizes insights from 257 papers to map current validation approaches and gaps.
- Presents cross-domain case studies to illustrate validation challenges.
Notable insights
- The paper identifies temporal validity and runtime evidence maintenance as underdeveloped areas in agentic AI validation.
- The use of cross-domain case studies provides practical insights into the recurring validation challenges in safety-critical settings.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2607.29405v1 Announce Type: new Abstract: Agentic AI systems act through multi-step trajectories that combine planning, tool use, memory, interaction, and adaptation. This behavior stretches validation practice beyond component testing and one-shot input--output evaluation, because acceptable system behavior now depends on how decisions unfold over time and under changing environmental conditions. This survey synthesizes 257 papers spanning agent evaluation, software assurance, cyber-physical systems, runtime monitoring, and regulatory guidance in order to characterize the validation problem for agentic systems. The review is organized around a five-dimension taxonomy covering behavioral, safety, temporal, regulatory, and multi-agent concerns, and uses that taxonomy to map current approaches and expose recurrent coverage gaps. The analysis shows that behavioral evaluation is comparatively mature, while temporal validity, runtime evidence maintenance, regulatory legibility, and open-ended multi-agent systems assurance remain under-developed. Three cross-domain case studies (medical care, industrial operations, smart-mobility systems) provide operational illustrations of how the five taxonomy dimensions recur in safety-critical settings, grounded in the failure patterns documented in the reviewed literature. The paper concludes with a lifecycle-oriented research agenda centered on bounded-autonomy specifications, adversarial trajectory generation, runtime monitoring, and audit-ready evidence structures. The central claim is that trustworthy deployment of agentic AI depends on validating trajectories in context rather than assessing isolated components alone.