arXiv AI

Evaluating Agentic Configuration Repair for Computer Networks

arXiv:2606. 06212v1 Announce Type: new Abstract: Misconfigurations in computer networks remain a major source of critical Internet outages.

arXiv AI
Sep 11

Can AI Agents Detect and Repair Artifact Drift in Network Experiments?

The paper introduces NetArtifactBench, a benchmark designed to evaluate whether AI agents can detect and repair inconsistencies in network experiment records while preserving supported claims. It tests 23 agent configurations on 52 instances with injected inconsistencies, finding an average pass rate of 65.3 % but no runtime exceeding 30 % for complex repairs that require recovering implicit relations and propagating changes across artifacts. The results highlight a clear distinction between local corrections and full record-level repair, leading the authors to argue that artifact integrity should be a primary design and evaluation criterion for AI agents in network systems.

By Tianzhu Zhang, Weichen Tao, Changgang Zheng, Yusheng Zheng, Long Chen, Xiaoyi Fan, Meikang Qiu
arXiv AI
Jul 16

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

arXiv:2607. 13705v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical.

By Zichen Ding, Jiaye Ge, Shufan Jiang, Kai Chen, Mo Li, Qingqiu Li, Zehao Li, Zonglin Li, Tiaohao Liang, Shudong Liu, Zerun Ma, Zixing Shang, Wenhui Tian, Zun Wang, Liwei Wu, Zhenyu Wu, Jun Xu, Bowen Yang, Dingbo Yuan, Qi Zhang, Songyang Zhang, Peiheng Zhou, Dongsheng Zhu
arXiv AI
Aug 13

CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations

arXiv:2608. 12002v1 Announce Type: new Abstract: Agents are increasingly considered for automating network operations and maintenance, where engineers must diagnose network faults, optimize configurations to enhance services, and reduce operational costs while acting under strict constraints.

By Xingyu Yan, Tingting Dai, Antonio De Domenico, Mohamed Sana, Nicola Piovesan, Changchang Li, Bowen Liu, Kun Jiang, Mengjie Zhang, Dingcheng Shan, Jing-Cheng Pang, Chenwei Wu, Sijie Wu, Lianying Chao, Haoran Cai, Jiantao Ye, Xubin Li, Simon Mark Lucas, Xin Chen
Hugging Face Trending Papers
Aug 27

FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets

FaulT-Bench is a new benchmark comprising 200 network troubleshooting scenarios across eight topologies, designed to test large‑language‑model agents on realistic, noisy tickets that may contain false premises or incorrect fault claims. The benchmark includes 72 rewritten tickets that vary reporter confidence and detail, and evaluates agents via an automated harness that scores diagnoses on outcome, fix, and reasoning quality. Results show that while agents perform well on accurate tickets, they degrade sharply on healthy networks with misleading reports, highlighting the importance of ticket wording over content.