arXiv AI

Incident-Arena: Getting agents to the last nine of reliability

Incident‑Arena is a new benchmark for AI coding agents focused on production incident response, featuring 20 tasks derived from real‑world open‑source software. Each task deploys a production application on an ephemerally created Kubernetes cluster, injects faults at various layers, and applies a sustained load profile. The benchmark introduces functional verifiers that maintain system‑level metrics while ensuring safe repairs, and shows that current frontier models achieve below 64.3% across the tasks, highlighting challenges in diagnosis, repair, and regression safety.

arXiv AI
Jun 30

A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis

arXiv:2606. 29193v1 Announce Type: cross Abstract: LLM-based agents are reshaping microservice operations into AgentOps, where benchmarks are key to evaluating failure diagnosis over multimodal observability data.

By Yuanhong Cai, Xiaohui Nie, Kanglin Yin, Changhua Pei, Yongqian Sun, Shenglin Zhang, Haibin Liu, Guiyang Liu, Xidao Wen, Fang Situ, Dan Pei
arXiv AI
Sep 17

ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software

ERPBench introduces a new evaluation paradigm for computer-use agents that operate via screenshots and simulated actions, focusing on enterprise software such as ERP systems. The benchmark tests agents on a live, reproducible ERP platform and scores tasks against ground-truth database values, highlighting challenges like dense interfaces, multi-step interactions, and persistent record errors. Experiments with six agents show that strong general GUI performance does not translate to reliable enterprise outcomes, with many agents frequently saving incorrect data.

By Kratika Bhagtani, Kusha Sridhar, Maziyar Baran Pouyan, Yuying Zhao, Eugene Siow
arXiv AI
Sep 1

AssetOpsBench: Benchmarking AI Agents for Task Automation in Industrial Asset Operations and Maintenance

arXiv:2506.03828v4 Announce Type: replace Abstract: AI for Industrial Asset Lifecycle Management aims to automate complex operational workflows, such as condition monitoring and maintenance schedulin...

By Dhaval Patel, Shuxin Lin, James Rayfield, Nianjun Zhou, Chathurangi Shyalika, Suryanarayana R Yarrabothula, Roman Vaculin, Natalia Martinez, Fearghal O'donncha, Jayant Kalagnanam
arXiv AI
3d ago

PANDA: A Decentralized Architecture with Flexible Orchestration for Scalable, Fault-Tolerant Multi-Agent Systems

PANDA is a decentralized architecture for large-scale, fault-tolerant multi-agent systems that enables heterogeneous agents to discover each other's capabilities and self-organize into specialized teams for each task. It decouples collective communication from team communication, allowing agents to participate in multiple teams simultaneously and load-balance tasks across the collective. PANDA supports three planning and execution patterns—star, chain, and mesh—detects and recovers from infrastructure and orchestration failures, and uses a web-of-trust model for governance without a central bottleneck.

By Matthew D. Laws, Cristina Nita-Rotaru