arXiv Machine Learning

TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution

TraceBench is a simulation-based framework that generates controlled root‑cause attribution tasks for time‑series data. In each task, an LLM agent must determine whether a system parameter was altered during a simulation of a physical dynamical system and identify the altered parameter. The authors evaluated four LLM agents on tasks derived from three interpretable mechanical systems, finding that agents perform better with domain context, rely mainly on numerical console output, and struggle more when required to produce Python scripts for labeling than when submitting direct predictions.

arXiv AI
Aug 7

TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories

arXiv:2608. 06346v1 Announce Type: new Abstract: LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging.

By Yunjia Qi, Zehua Yin, Xintong Shi, Hao Peng, Songyuanyi Lu, Yixian Liu, Richeng Xuan, Yuhong Liu, Zhichao Hu, Xiaozhi Wang, Lei Hou, Bin Xu, Juanzi Li
arXiv AI
Sep 15

Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures

The paper introduces Continual Search, an iterative framework that guides large language models to persistently search for diagnostic evidence in long AI agent execution logs, addressing the limitations of one-shot judgments. Evaluated on four existing RCA benchmarks and a new large-scale dataset called MegaRCA-Mix, Continual Search consistently boosts attribution performance, achieving a 40% F1 improvement for GPT‑5.5 on MegaRCA‑Mix. The results show that effective search can outweigh raw model scale, enabling lower-tier models to outperform higher-tier ones in root‑cause attribution tasks.

By Harsh Raj, David Lee, Anas Mahmoud, Renxiong Wang, Razvan-Gabriel Dumitru, Chenguang Wang, Tong Zhao, Yunzhong He, Darvin Yi, Vipul Gupta
arXiv AI
Jun 8

TRACE: Trajectory Reasoning through Adaptive Cross-Step Evidence Aggregation for LLM Agents

arXiv:2606. 07054v1 Announce Type: cross Abstract: Autonomous LLM agents can pursue hidden malicious objectives through sequences of individually benign actions, making sabotage difficult to detect using standard trajectory-level monitoring.

By Vijitha Mittapalli, Shreyaa Jayant Dani, Satya Srujana Pilli, Snigdha Ansu, Mohammadreza Teymoorianfard, Franck Dernoncourt, Hongjie Chen, Yu Wang, Ryan A. Rossi, Nesreen K. Ahmed
arXiv AI
Sep 24

TimeEvo: Failure-Driven Self-Evolution of a Time Series Agent

TimeEvo is a new method for time‑series agents that autonomously evolves its tool library based on failures observed during runtime. By clustering diagnosed failures into capability gaps, planning measurements, synthesizing evidence‑only tools, and admitting candidates through a paired gate, the system starts from an empty library and improves accuracy across ten QA tasks and three backbones. Experiments show that even a library built on a cheap model benefits stronger models when installed.

By Jie Yang, Yan Zheng, Jiarui Sun, Xiran Fan, Junpeng Wang, Liang Wang, Zelin Xu, Qinghua Liu, Zhengyu Fang, Yiwei Cai, Philip S. Yu
arXiv AI
4d ago

MAADBench: The Refreshable Paradigm for Anomaly Detection in Multi-Agent Systems

MAADBench is a refreshable benchmark for anomaly detection in multi‑agent systems powered by large language models. It addresses the challenge of keeping benchmarks current by sampling and coupling generative tasks, generating trace data under configurable LLM backbones, and automatically providing deterministic step‑level labels. The authors evaluated 25 anomaly‑detection methods on 5,200 labeled traces, finding that existing approaches depend heavily on supervision, struggle with subtle MAS‑specific anomalies, and lack robustness across different LLM backbones.

By Lei Ma, Dennis Hofmann, Haowen Xu, Joshua DeOliveira, Peter VanNostrand, Lei Cao, Elke Rundensteiner
arXiv AI
6d ago

Detecting Time Series Anomalies Like an Expert: A Multi-Agent LLM Framework with Specialized Analyzers

The paper introduces SAGE, a multi‑agent framework that uses specialized analyzers to diagnose univariate time‑series anomalies by examining point, structural, seasonal, and pattern deviations. Each analyzer produces numerical evidence and visual diagnostics, which a Detector consolidates into intervals, candidate types, and confidence scores, and a Supervisor converts these into analyst‑friendly reports. Experiments on Yahoo S5, KPI, and WSD datasets show SAGE achieving the highest average Point‑F1 score (66.26) and receiving higher usefulness ratings in a blind human study.

By Hyeongwon Kang, Jeongseob Kim, Jinwoo Park, Pilsung Kang