arXiv AI By Hyeongwon Kang, Jeongseob Kim, Jinwoo Park, Pilsung Kang

Detecting Time Series Anomalies Like an Expert: A Multi-Agent LLM Framework with Specialized Analyzers

Read the original on arXiv AI →

The paper introduces SAGE, a multi‑agent framework that uses specialized analyzers to diagnose univariate time‑series anomalies by examining point, structural, seasonal, and pattern deviations. Each analyzer produces numerical evidence and visual diagnostics, which a Detector consolidates into intervals, candidate types, and confidence scores, and a Supervisor converts these into analyst‑friendly reports. Experiments on Yahoo S5, KPI, and WSD datasets show SAGE achieving the highest average Point‑F1 score (66.26) and receiving higher usefulness ratings in a blind human study.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

Structured Frequency-Domain Evidence for LLM-Based Time-Series Anomaly Detection

The paper introduces a zero‑shot time‑series anomaly detection framework that augments traditional indexed, de‑seasonalized observations with compact frequency‑domain evidence derived from the Fast Fourier Transform. This evidence is provided at two resolutions: a global summary of sequence‑level periodicity and a local snapshot of time‑localized spectral deviations. Experiments on the AnomLLM benchmark using several large language models—including InternVL2‑LLaMA3‑76B, Qwen2.5‑VL‑72B‑Instruct, Gemini‑2.5‑Flash, and GPT‑4o—demonstrate that incorporating explicit frequency‑domain evidence improves anomaly detection performance over existing LLM‑based baselines.

By Jungwook Seo, Sangwon Son, Minjeong Kim, Seungmin Han, Seojin Yoo, Sungyong Baik
arXiv Machine Learning
Aug 28

TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution

TraceBench is a simulation-based framework that generates controlled root‑cause attribution tasks for time‑series data. In each task, an LLM agent must determine whether a system parameter was altered during a simulation of a physical dynamical system and identify the altered parameter. The authors evaluated four LLM agents on tasks derived from three interpretable mechanical systems, finding that agents perform better with domain context, rely mainly on numerical console output, and struggle more when required to produce Python scripts for labeling than when submitting direct predictions.

By Tommaso Bendinelli, Artur Dox, Christian Holz