arXiv AI

Beyond Component Testing: Validating Agentic AI Systems

arXiv:2607. 29405v1 Announce Type: new Abstract: Agentic AI systems act through multi-step trajectories that combine planning, tool use, memory, interaction, and adaptation.

arXiv AI
Jul 22

Engineering Trustworthy Agentic AI for Critical Systems

arXiv:2607. 18548v1 Announce Type: new Abstract: Agentic artificial intelligence systems, capable of autonomous perception, planning, tool use, and multi-step action, are increasingly proposed for critical engineering domains where decisions carry physical, operational, or economic consequences.

By Omar Al-Refai, Ibrahim Shahbaz, Adam Ali Husseinat, Michael Mandulak, Jaewon Kim, Eman Hammad
arXiv AI
Sep 7

From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments

The paper reviews how large language models have evolved into agents that can influence external environments through tool use, interface operation, delegation, state retention, virtual world inhabitation, and robotic control. It critiques the narrative of a single march toward autonomy, distinguishing model competence from system integration, persistence, and safe authority. The authors find that action-interface expansion is well documented, while robust completion, recovery, authorization, and independent verification remain less proven, and they propose a framework of justified delegation to guide future research.

By Linsen Zhu, Mengqing Cai
arXiv AI
Sep 25

Where Cyber Agents Struggle: Bottleneck Analysis of Multi-Stage LLM Agents

The paper presents a diagnostic study of a multi‑stage LLM‑based cyber agent system, examining its orchestrator, executor, and validator components in enterprise‑style lateral‑movement scenarios. Six advanced LLMs were tested across expert‑defined, self‑scaffolded, and fully autonomous modes, with metrics that include validator consistency, evidence grounding, token usage, retries, and runtime. Findings show that while validators are generally relevant, they are often nonspecific and overly optimistic, and the main bottlenecks lie in credential acquisition and lateral‑movement tasks, especially under full autonomy.

By Saeedeh Lohrasbi, Mohammad Mamun, Ahmed Yehia, Scott Buffett, Sherif Saad
arXiv AI
Jun 17

Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety

arXiv:2605. 12729v2 Announce Type: replace-cross Abstract: Large language models are increasingly being used to support network operations (NetOps) and artificial intelligence for IT operations (AIOps), including incident investigation, root-cause analysis, configuration synthesis, and limited self-healing.

By Muhammad Bilal, Jon Crowcroft, Ruizhi Wang, Xiaolong Xu, Schahram Dustdar
arXiv AI
Aug 24

Testing and Evaluation of Agentic AI Systems In Military Command and Control

The paper examines how agentic AI systems intended for military command and control are tested and evaluated. It reviews 240 testing practices across eight dimensions and three lifecycle stages, uncovering eight assumptions—grouped into system specifiability, stability, composability, and supervisability—whose validity is weakened by agentic properties. Consequently, test results may meet procedural standards but do not guarantee that fielded behavior matches tested behavior, leading the authors to propose ten assurance claims and suggest that uncertainty be managed through deployment‑time monitoring and defined expiry conditions.

By Ulysse Richard, Heather Frase, Sarah Cao, Di Cooke, Sebastian Kwon, Adrianna Tan