arXiv AI

Detecting Time Series Anomalies Like an Expert: A Multi-Agent LLM Framework with Specialized Analyzers

The paper introduces SAGE, a multi‑agent framework that uses specialized analyzers to diagnose univariate time‑series anomalies by examining point, structural, seasonal, and pattern deviations. Each analyzer produces numerical evidence and visual diagnostics, which a Detector consolidates into intervals, candidate types, and confidence scores, and a Supervisor converts these into analyst‑friendly reports. Experiments on Yahoo S5, KPI, and WSD datasets show SAGE achieving the highest average Point‑F1 score (66.26) and receiving higher usefulness ratings in a blind human study.

arXiv AI
Aug 26

Structured Frequency-Domain Evidence for LLM-Based Time-Series Anomaly Detection

The paper introduces a zero‑shot time‑series anomaly detection framework that augments traditional indexed, de‑seasonalized observations with compact frequency‑domain evidence derived from the Fast Fourier Transform. This evidence is provided at two resolutions: a global summary of sequence‑level periodicity and a local snapshot of time‑localized spectral deviations. Experiments on the AnomLLM benchmark using several large language models—including InternVL2‑LLaMA3‑76B, Qwen2.5‑VL‑72B‑Instruct, Gemini‑2.5‑Flash, and GPT‑4o—demonstrate that incorporating explicit frequency‑domain evidence improves anomaly detection performance over existing LLM‑based baselines.

By Jungwook Seo, Sangwon Son, Minjeong Kim, Seungmin Han, Seojin Yoo, Sungyong Baik
arXiv Machine Learning
Aug 28

TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution

TraceBench is a simulation-based framework that generates controlled root‑cause attribution tasks for time‑series data. In each task, an LLM agent must determine whether a system parameter was altered during a simulation of a physical dynamical system and identify the altered parameter. The authors evaluated four LLM agents on tasks derived from three interpretable mechanical systems, finding that agents perform better with domain context, rely mainly on numerical console output, and struggle more when required to produce Python scripts for labeling than when submitting direct predictions.

By Tommaso Bendinelli, Artur Dox, Christian Holz
arXiv AI
Jun 2

AnomSeer: Reinforcing Multimodal LLMs to Reason for Time-Series Anomaly Detection

arXiv:2602. 08868v2 Announce Type: replace-cross Abstract: Time-series anomaly detection (TSAD) with multimodal large language models (MLLMs) is an emerging area, yet a persistent challenge remains: MLLMs rely on coarse time-series heuristics but struggle with multi-dimensional, detailed reasoning, which is vital for understanding complex time-series data.

By Junru Zhang, Lang Feng, Haoran Shi, Xu Guo, Han Yu, Yabo Dong, Duanqing Xu
arXiv AI
Jun 8

TSAQA: Time Series Analysis Question And Answering Benchmark

arXiv:2601. 23204v2 Announce Type: replace Abstract: Time series data are integral to critical applications across domains such as finance, healthcare, transportation, and environmental science.

By Baoyu Jing, Sanhorn Chen, Lecheng Zheng, Boyu Liu, Zihao Li, Jiaru Zou, Tianxin Wei, Zhining Liu, Zhichen Zeng, Ruizhong Qiu, Xiao Lin, Yuchen Yan, Dongqi Fu, Jingchao Ni, Jingrui He, Hanghang Tong