arXiv AI By Yongzi Yu, Ao Li, Le Wang, Ziyue Li, Fugee Tsung, Yuxuan Liang, Man Li

Plan First, Judge Later, Run Better: A DMAIC-Inspired Agentic System for Industrial Anomaly Detection

Read the original on arXiv AI →

arXiv:2606. 04599v1 Announce Type: new Abstract: Large language model (LLM) agents have shown promise in automating complex data-analysis workflows, but their reliable deployment remains challenging in high-stakes industrial scenarios.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
4d ago

MAADBench: The Refreshable Paradigm for Anomaly Detection in Multi-Agent Systems

MAADBench is a refreshable benchmark for anomaly detection in multi‑agent systems powered by large language models. It addresses the challenge of keeping benchmarks current by sampling and coupling generative tasks, generating trace data under configurable LLM backbones, and automatically providing deterministic step‑level labels. The authors evaluated 25 anomaly‑detection methods on 5,200 labeled traces, finding that existing approaches depend heavily on supervision, struggle with subtle MAS‑specific anomalies, and lack robustness across different LLM backbones.

By Lei Ma, Dennis Hofmann, Haowen Xu, Joshua DeOliveira, Peter VanNostrand, Lei Cao, Elke Rundensteiner