arXiv AI

Autonomous Incident Resolution at Hyperscale: An Agentic AI Architecture for Network Operations

arXiv:2606. 09122v1 Announce Type: cross Abstract: Cloud network infrastructure at hyperscale presents unique operational challenges where traditional human-driven incident response cannot keep pace with the volume, velocity, and complexity of failures.

arXiv AI
Jul 7

AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression

arXiv:2512. 13956v4 Announce Type: replace-cross Abstract: Cloud-native systems have made operational work both more powerful and harder to automate: incidents unfold across microservices, logs and metrics arrive faster than operators can inspect them, and recovery actions must be coordinated without losing the causal context that makes them safe.

By Zishan Bai, Hanxuan Chen, Jiayi Gu, Wenqian Weng, Enze Ge, Jiacheng Shi, Yichao Zhang, Zhimo Han, Riyang Bao, Xinyuan Song, Jacqueline Pang, Junfeng Hao
arXiv AI
Aug 3

OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems

arXiv:2607. 28629v1 Announce Type: new Abstract: The rapid transition from reactive large language models (LLMs) to persistent, action-capable systems has exposed critical gaps in the architectural understanding of Agentic AI, particularly in separating inference, orchestration, and execution layers for autonomous AI agents.

By Konstantinos I. Roumeliotis, Ranjan Sapkota
arXiv AI
Jul 28

Building AI That Works: ESnet's Pragmatic Approach to AI-Driven Operational Excellence

arXiv:2607. 22948v1 Announce Type: cross Abstract: The ORBIT (Operations Responses and Business Intelligence Toolkit) project was initiated to assess agentic AI for the upcoming ESnet 7 initiative and to address persistent operational pain points in the Network Operations Center (NOC) workflow.

By Bin Dong, Sukhada Gholba, Brooklin Gore, Shawn Kwang, David Mitchell, Samuel Oehlert, Garrett Stewart, Brendan White, Luke Baker, Ed Balas, Britt Gathright, Chin Guok, Jon-Paul Heron, John MacAuley, Scott Richmond, Chris Robb, Chris Tracy, Kesheng Wu
arXiv AI
2d ago

Incident-Arena: Getting agents to the last nine of reliability

Incident‑Arena is a new benchmark for AI coding agents focused on production incident response, featuring 20 tasks derived from real‑world open‑source software. Each task deploys a production application on an ephemerally created Kubernetes cluster, injects faults at various layers, and applies a sustained load profile. The benchmark introduces functional verifiers that maintain system‑level metrics while ensuring safe repairs, and shows that current frontier models achieve below 64.3% across the tasks, highlighting challenges in diagnosis, repair, and regression safety.

By Andre Fu, Malik Drabla, Leon Liu, Meji Abidoye, Marek Suppa, Lata Mishra, Adnan El Assadi, Yiyuan Li
arXiv AI
Sep 1

Forward-Deployed Full-Stack Engineering for Autonomous Cloud MLOps

The paper introduces a multi‑agent framework that transforms natural‑language MLOps tasks into verified repositories and operational cloud deployments. It uses a stateful Graph Orchestrator to coordinate agents for repository generation, review, execution, verification, release, and monitoring, ensuring lifecycle transitions only occur when supported by verifiable evidence. The framework, implemented on Google Cloud Platform, demonstrates prevention of unsupported transitions and drives each run toward a verified deployment or an auditable failure.

By Sagar Srinivas Sakhinana, Venkataramana Runkana
arXiv AI
3d ago

PANDA: A Decentralized Architecture with Flexible Orchestration for Scalable, Fault-Tolerant Multi-Agent Systems

PANDA is a decentralized architecture for large-scale, fault-tolerant multi-agent systems that enables heterogeneous agents to discover each other's capabilities and self-organize into specialized teams for each task. It decouples collective communication from team communication, allowing agents to participate in multiple teams simultaneously and load-balance tasks across the collective. PANDA supports three planning and execution patterns—star, chain, and mesh—detects and recovers from infrastructure and orchestration failures, and uses a web-of-trust model for governance without a central bottleneck.

By Matthew D. Laws, Cristina Nita-Rotaru