arXiv Machine Learning

Agentic Root Cause Analysis through Evidence-Grounded Reasoning

arXiv:2607. 22385v1 Announce Type: cross Abstract: Diagnosing the root cause of anomalies is essential for safe industrial operation.

Hugging Face Trending Papers
Jul 15

How Far Can Root Cause Analysis Go on Real-World Telemetry Data?

Identifying root causes in production microservice failures requires reasoning over large-scale, multimodal telemetry spanning metrics, logs, and traces, a problem that has proved resistant to both classical and LLM-based approaches. The OpenRCA dataset exemplifies these challenges: it is large-scale, multimodal, and lacks detailed domain knowledge, and yields consistently low accuracy across all existing methods.

arXiv AI
Jun 30

A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis

arXiv:2606. 29193v1 Announce Type: cross Abstract: LLM-based agents are reshaping microservice operations into AgentOps, where benchmarks are key to evaluating failure diagnosis over multimodal observability data.

By Yuanhong Cai, Xiaohui Nie, Kanglin Yin, Changhua Pei, Yongqian Sun, Shenglin Zhang, Haibin Liu, Guiyang Liu, Xidao Wen, Fang Situ, Dan Pei
Hugging Face Trending Papers
Jun 25

OpenRCA 2.0: From Outcome Labels to Causal Process Supervision

Root cause analysis (RCA) poses a holistic test of LLM agentic capabilities, such as long-context understanding, multi-step reasoning, and tool use. However, existing datasets suffer from a fundamental gap: they label only the root cause, not the propagation path connecting it to the observed symptom, which largely simplifies the task to naive pattern matching.

arXiv AI
Aug 28

Learning to Predict, Discover, and Reason in High-Dimensional Event Sequences

The paper proposes a new framework for automated fault diagnostics in modern vehicles by treating diagnostic trouble codes (DTCs) as a high‑dimensional language. It introduces Transformer‑based models for predictive maintenance, scalable causal discovery methods, and a multi‑agent system that automatically generates Boolean error‑pattern rules. The approach aims to replace costly manual grouping of DTCs with scalable, data‑driven techniques.

By Hugo Math
Hugging Face Trending Papers
Sep 2

Large Language Models (LLMs) for Telecom Root Cause Analysis (RCA): A Structured Reasoning Framework for Evidence-Grounded Diagnosis

The paper discusses the challenges of root cause analysis (RCA) in 5G and 6G telecom networks, where complex cross-layer dependencies make diagnosis difficult. It reviews the progression from rule‑based and machine‑learning RCA methods to emerging large language model (LLM) approaches, highlighting issues such as hallucination and unstable reasoning when using vanilla LLMs. The authors propose a structured reasoning framework that organizes network telemetry into canonical contexts, enforces decision‑path reasoning, and generates evidence‑grounded explanations, showing improved diagnostic accuracy on two 5G RCA datasets.

arXiv AI
Aug 18

LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures

arXiv:2608. 15242v1 Announce Type: new Abstract: When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory.

By Yunfei Zhang, Boyu Feng, Changhua Pei, Zexin Wang, Zhihuang Peng, Xinlong Liu, Hengyue Jiang, Difeng Ma, Jiayi Zhang, Yongzhou Yao, Yanan Zhao, Fei Sun, Yintong Huo, Zhaoyang Liu, Jingjing Li, Gaogang Xie, Dan Pei
arXiv Machine Learning
Aug 28

TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution

TraceBench is a simulation-based framework that generates controlled root‑cause attribution tasks for time‑series data. In each task, an LLM agent must determine whether a system parameter was altered during a simulation of a physical dynamical system and identify the altered parameter. The authors evaluated four LLM agents on tasks derived from three interpretable mechanical systems, finding that agents perform better with domain context, rely mainly on numerical console output, and struggle more when required to produce Python scripts for labeling than when submitting direct predictions.

By Tommaso Bendinelli, Artur Dox, Christian Holz