arXiv AI By Yufei Xia, Anjun Gao, Yueyang Quan, Zhuqing Liu, Minghong Fang

Who Broke the System? Failure Localization in LLM-Based Multi-Agent Systems

Read the original on arXiv AI →

arXiv:2607. 07989v1 Announce Type: cross Abstract: Large language model (LLM) based multi-agent systems enable complex problem solving through coordinated reasoning and action, but their distributed structure also introduces new challenges in diagnosing system-level failures.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 2

POIROT: Interrogating Agents for Failure Detection in Multi-Agent Systems

arXiv:2606. 02282v1 Announce Type: new Abstract: Orchestrating Large Language Models into Multi-Agent Systems (LLM-MAS) has unlocked remarkable reasoning capabilities, yet emergent failures and hallucinations that resist characterisation block their deployment in safety-critical domains -- a gap made legally untenable by emerging AI regulation.

By I\~naki Dellibarda Varela, R. Sendra-Arranz, Pablo Romero-Sorozabal, J. M. Valverde-Garc\'ia, Annemarie F. Laudanski, \'Alvaro Guti\'errez, Eduardo Rocon, Manuel Cebrian
arXiv AI
Sep 3

Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral Abstractions

The paper introduces AGENTSCOPE, a neuro‑symbolic method for diagnosing failures in large language model agents. It abstracts agent trajectories into structured representations and employs neural invariants to define behavior properties. Using LLM‑guided reasoning on these abstractions, AGENTSCOPE identifies both the failure step and its type, outperforming existing techniques on several datasets.

By Jiayi Bi, Yanjie Gao, Yuanmin Xie, Liqun Li, Tianyin Xu, Fan Yang, Mao Yang