arXiv:2609.06835v1 Announce Type: cross
Abstract: Agentic AI systems execute complex tasks through long-horizon workflows of planning, tool use, and multi-agent coordination. Task failures in these s...
By Chaoyu Zhang, Hexuan Yu, Heng Jin, Shanghao Shi, Ning Zhang, Yi Shi, Yulia R. Gel, Y. Thomas Hou, Wenjing Lou
MAADBench is a refreshable benchmark for anomaly detection in multi‑agent systems powered by large language models. It addresses the challenge of keeping benchmarks current by sampling and coupling generative tasks, generating trace data under configurable LLM backbones, and automatically providing deterministic step‑level labels. The authors evaluated 25 anomaly‑detection methods on 5,200 labeled traces, finding that existing approaches depend heavily on supervision, struggle with subtle MAS‑specific anomalies, and lack robustness across different LLM backbones.
By Lei Ma, Dennis Hofmann, Haowen Xu, Joshua DeOliveira, Peter VanNostrand, Lei Cao, Elke Rundensteiner
arXiv:2602. 13807v2 Announce Type: replace Abstract: Time series anomaly detection is critical in many real-world applications, where effective solutions must localize anomalous regions and support reliable decision-making under complex settings.
By Xiaoyu Tao, Yuchong Wu, Mingyue Cheng, Ze Guo, Tian Gao
arXiv:2607. 07052v1 Announce Type: cross Abstract: AI agents deployed for IT operations are typically permanent cost centers because every execution requires full LLM inference, even for previously solved problems.
By Arun Malik
Production LLM agents often waste latency and reliability by regenerating code for the same procedural steps on every request. We replace this inference-time coding loop with an agentic tool-making pipeline that compiles repeated SOP steps into validated, versioned tools before deployment.
arXiv:2607. 08010v1 Announce Type: cross Abstract: Production LLM agents often waste latency and reliability by regenerating code for the same procedural steps on every request.
By Kalle Kujanp\"a\"a, Ning Liu, Shahnawaz Alam, Yeshwanth Reddy Sura, Tianyu Yang, Kristina Klinkner, Shervin Malmasi
arXiv:2608. 11679v1 Announce Type: new Abstract: Digital twins are increasingly used to monitor and simulate the behavior of cyber-physical systems.
By Touseef Hasan, Mounika Ghanta, Souvika Sarkar, Ujjwal Guin
arXiv:2606. 02282v1 Announce Type: new Abstract: Orchestrating Large Language Models into Multi-Agent Systems (LLM-MAS) has unlocked remarkable reasoning capabilities, yet emergent failures and hallucinations that resist characterisation block their deployment in safety-critical domains -- a gap made legally untenable by emerging AI regulation.
By I\~naki Dellibarda Varela, R. Sendra-Arranz, Pablo Romero-Sorozabal, J. M. Valverde-Garc\'ia, Annemarie F. Laudanski, \'Alvaro Guti\'errez, Eduardo Rocon, Manuel Cebrian
arXiv:2606. 05806v1 Announce Type: new Abstract: Existing benchmarks evaluate Tool-Integrated Reasoning (TIR) in LLMs on idealized ''happy paths'', largely overlooking real-world tool failures.
By Dongsheng Zhu, Xuchen Ma, Yucheng Shen, Xiang Li, Yukun Zhao, Shuaiqiang Wang, Lingyong Yan, Dawei Yin
The survey reviews LLM-based agents tailored to time-series tasks, organizing them by problem type rather than technical components. It categorizes existing systems into forecasting and reasoning, augmentation and synthesis, anomaly detection and diagnosis, and decision support, and analyzes how task demands influence agent architecture, tool use, and memory design. The paper also summarizes datasets, environments, and compares model performance, providing a task-oriented guide and highlighting gaps for future research.
By Yilong Chen, Xiao Qin, Chenghao Liu, Liang Wu, Noelle I. Samia, Kaize Ding
arXiv:2508.15526v2 Announce Type: replace
Abstract: The rapid proliferation of large language models (LLMs) has intensified the requirement for reliable safety evaluation to uncover model vulnerabili...
By Xiangyang Zhu, Yuan Tian, Chunyi Li, Kaiwei Zhang, Wei Sun, Guangtao Zhai
arXiv:2607. 18142v1 Announce Type: cross Abstract: Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, which is crucial for modern manufacturing and quality control systems.
By Mei Yuan, Qi Long, Qifeng Wu, Zhenyang Li, Yizhou Zhao, Lei Wang, Yang Liu, Min Xu