arXiv:2608. 12002v1 Announce Type: new Abstract: Agents are increasingly considered for automating network operations and maintenance, where engineers must diagnose network faults, optimize configurations to enhance services, and reduce operational costs while acting under strict constraints.
By Xingyu Yan, Tingting Dai, Antonio De Domenico, Mohamed Sana, Nicola Piovesan, Changchang Li, Bowen Liu, Kun Jiang, Mengjie Zhang, Dingcheng Shan, Jing-Cheng Pang, Chenwei Wu, Sijie Wu, Lianying Chao, Haoran Cai, Jiantao Ye, Xubin Li, Simon Mark Lucas, Xin Chen
arXiv:2607. 22948v1 Announce Type: cross Abstract: The ORBIT (Operations Responses and Business Intelligence Toolkit) project was initiated to assess agentic AI for the upcoming ESnet 7 initiative and to address persistent operational pain points in the Network Operations Center (NOC) workflow.
By Bin Dong, Sukhada Gholba, Brooklin Gore, Shawn Kwang, David Mitchell, Samuel Oehlert, Garrett Stewart, Brendan White, Luke Baker, Ed Balas, Britt Gathright, Chin Guok, Jon-Paul Heron, John MacAuley, Scott Richmond, Chris Robb, Chris Tracy, Kesheng Wu
FaulT-Bench is a new benchmark comprising 200 network troubleshooting scenarios across eight topologies, designed to test large‑language‑model agents on realistic, noisy tickets that may contain false premises or incorrect fault claims. The benchmark includes 72 rewritten tickets that vary reporter confidence and detail, and evaluates agents via an automated harness that scores diagnoses on outcome, fix, and reasoning quality. Results show that while agents perform well on accurate tickets, they degrade sharply on healthy networks with misleading reports, highlighting the importance of ticket wording over content.
arXiv:2606. 06212v1 Announce Type: new Abstract: Misconfigurations in computer networks remain a major source of critical Internet outages.
By Rufat Asadli, Benjamin Hoffman, Ioannis Protogeros, Laurent Vanbever
arXiv:2606. 29116v1 Announce Type: new Abstract: Large Language Models (LLMs) are rapidly being adopted in low-code and no-code automation platforms, where non-expert users design workflows that combine natural language understanding with external services and APIs.
By Yutian Tang, Yuming Zhou, Huaming Chen
arXiv:2605. 12729v2 Announce Type: replace-cross Abstract: Large language models are increasingly being used to support network operations (NetOps) and artificial intelligence for IT operations (AIOps), including incident investigation, root-cause analysis, configuration synthesis, and limited self-healing.
By Muhammad Bilal, Jon Crowcroft, Ruizhi Wang, Xiaolong Xu, Schahram Dustdar
arXiv:2608.23179v1 Announce Type: cross
Abstract: Large language model (LLM) agents are increasingly attractive for automating network configuration, yet their reliability and failure patterns are po...
By Chang Liu, Xiaohui Xie, Xinyi Chen, Yong Cui
arXiv:2512. 22256v2 Announce Type: replace-cross Abstract: Software issue resolution aims to address real-world issues in software repositories based on natural language descriptions provided by users, and represents a key aspect of software maintenance.
By Zhonghao Jiang, David Lo, Zhongxin Liu
arXiv:2609.06835v1 Announce Type: cross
Abstract: Agentic AI systems execute complex tasks through long-horizon workflows of planning, tool use, and multi-agent coordination. Task failures in these s...
By Chaoyu Zhang, Hexuan Yu, Heng Jin, Shanghao Shi, Ning Zhang, Yi Shi, Yulia R. Gel, Y. Thomas Hou, Wenjing Lou
arXiv:2602.02475v2 Announce Type: replace
Abstract: AI agents often fail in ways that are difficult to localize because executions are probabilistic, long-horizon, multi-agent, and mediated by noisy...
By Shraddha Barke, Arnav Goyal, Alind Khare, Avaljot Singh, Suman Nath, Chetan Bansal
ParaRecover is a new process-level benchmark designed to evaluate error localization and recovery in multi-turn parallel tool-use agents. It contains 10,626 instances across two difficulty levels, built on a taxonomy of 14 error types that cover planning dependencies, tool selection, and argument matching. The benchmark introduces the SDE rubric, which assesses structural integrity, diagnostic reasoning, and evolutionary strategy during agent execution, and demonstrates that it can guide improvements in agents’ reflective recovery capabilities.
By Bowen Guan, Zhentao Yin, Yanming Shen
FaulT-Bench is a new benchmark comprising 200 network troubleshooting scenarios across eight topologies, designed to test large‑language‑model agents on realistic, noisy user tickets that may contain false premises or incorrect fault claims. The benchmark includes 72 rewritten tickets that vary reporter confidence and detail while keeping the network state constant, allowing isolation of the impact of ticket wording on diagnosis. Evaluation of agents such as SADE, ReAct, and Claude Code shows they perform well on accurate tickets but degrade sharply on misleading or healthy‑network tickets, revealing differing failure modes and highlighting the importance of robust reasoning over unreliable input.
By Kuan-Hao Tseng, Niruth Bogahawatta, Yasod Ginige, Kunjan Patel, Kosta Dakic, Suranga Seneviratne