arXiv AI

MAADBench: The Refreshable Paradigm for Anomaly Detection in Multi-Agent Systems

MAADBench is a refreshable benchmark for anomaly detection in multi‑agent systems powered by large language models. It addresses the challenge of keeping benchmarks current by sampling and coupling generative tasks, generating trace data under configurable LLM backbones, and automatically providing deterministic step‑level labels. The authors evaluated 25 anomaly‑detection methods on 5,200 labeled traces, finding that existing approaches depend heavily on supervision, struggle with subtle MAS‑specific anomalies, and lack robustness across different LLM backbones.

arXiv AI
Jun 8

TRACE: Trajectory Reasoning through Adaptive Cross-Step Evidence Aggregation for LLM Agents

arXiv:2606. 07054v1 Announce Type: cross Abstract: Autonomous LLM agents can pursue hidden malicious objectives through sequences of individually benign actions, making sabotage difficult to detect using standard trajectory-level monitoring.

By Vijitha Mittapalli, Shreyaa Jayant Dani, Satya Srujana Pilli, Snigdha Ansu, Mohammadreza Teymoorianfard, Franck Dernoncourt, Hongjie Chen, Yu Wang, Ryan A. Rossi, Nesreen K. Ahmed
arXiv AI
Aug 7

TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories

arXiv:2608. 06346v1 Announce Type: new Abstract: LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging.

By Yunjia Qi, Zehua Yin, Xintong Shi, Hao Peng, Songyuanyi Lu, Yixian Liu, Richeng Xuan, Yuhong Liu, Zhichao Hu, Xiaozhi Wang, Lei Hou, Bin Xu, Juanzi Li
arXiv AI
Jul 31

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

arXiv:2607. 26791v1 Announce Type: cross Abstract: Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities.

By Lehan Wang, Boli Chen, Ruixue Ding, Pengjun Xie, Jinwei Huang, Zhendong Liu, Shuo Wang, Tao Lei, Xin Ouyang, Xiaomeng Li
arXiv Machine Learning
Sep 3

TrajMind: Chaining Role-Specialized LoRAs for Fast-and-Slow Collective Trajectory Anomaly Diagnosis

TrajMind is a framework for diagnosing collective anomalies in urban trajectory data. It separates continuous screening from on-demand diagnosis, using a fast text-only path for alerts and a slow vision‑language path that chains role‑specialized LoRA adapters for detailed, evidence‑backed what‑who‑where‑when records. Experiments show the slow path outperforms baselines by over 15 percentage points in typing and 13 in localization, while the fast path cuts latency by 41% and retains high accuracy.

By Jiahao Wu, Zhenqun Yang, Chen Jason Zhang, Qing Li
arXiv AI
Jul 14

AgentAbstain: Do LLM Agents Know When Not to Act?

arXiv:2607. 10059v1 Announce Type: new Abstract: Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain.

By Xun Liu, Yi Evie Zhang, Vira Kasprova, Parisa Rabbani, Pardis Sadat Zahraei, Tianyu Zhang, Ali Ebrahimpour-Boroojeny, Varun Chandrasekaran