arXiv AI By Tobias Labarta, Frederik Pahde, Novak Boskov, Maximilian Dreyer, David Birkenberger, Manzoor Ahmed Khan, Sebastian Lapuschkin, Wojciech Samek

Safety Signals to Verify NetOps Agents with Action-Level Granularity

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Sep 11

Can AI Agents Deliver Verifiable Network-Wide Outcomes Across Authority Boundaries?

The paper introduces EvidenceNet, a runtime assurance layer designed to verify that coordinated AI agent operations achieve an operator’s intended network-wide outcomes across multiple administrative domains. EvidenceNet collects post-change observations from the required authority scopes, checks their freshness and validity, and uses a verifier agent to assess observation content. Experiments on live routing networks demonstrate that this approach can detect successful outcomes that configuration-action logs alone miss, and it rejects completions when observations are sourced incorrectly, substituted, or stale.

By Tianzhu Zhang, Chih-Kai Huang, Meikang Qiu
arXiv AI
Sep 11

Can AI Agents Detect and Repair Artifact Drift in Network Experiments?

The paper introduces NetArtifactBench, a benchmark designed to evaluate whether AI agents can detect and repair inconsistencies in network experiment records while preserving supported claims. It tests 23 agent configurations on 52 instances with injected inconsistencies, finding an average pass rate of 65.3 % but no runtime exceeding 30 % for complex repairs that require recovering implicit relations and propagating changes across artifacts. The results highlight a clear distinction between local corrections and full record-level repair, leading the authors to argue that artifact integrity should be a primary design and evaluation criterion for AI agents in network systems.

By Tianzhu Zhang, Weichen Tao, Changgang Zheng, Yusheng Zheng, Long Chen, Xiaoyi Fan, Meikang Qiu
arXiv AI
Jun 17

Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety

arXiv:2605. 12729v2 Announce Type: replace-cross Abstract: Large language models are increasingly being used to support network operations (NetOps) and artificial intelligence for IT operations (AIOps), including incident investigation, root-cause analysis, configuration synthesis, and limited self-healing.

By Muhammad Bilal, Jon Crowcroft, Ruizhi Wang, Xiaolong Xu, Schahram Dustdar