arXiv AI

BRA-Audit: Budgeted Runtime Auditing for LLM Multi-Agent Systems via Cumulative-Exposure Audit-Point Placement

arXiv:2608. 14668v1 Announce Type: cross Abstract: LLM-based multi-agent systems (LLM-MAS) solve complex tasks through specialized collaboration, but inter-agent dependencies can propagate hallucinated or malicious outputs into system-level failures.

arXiv AI
6d ago

AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents

AgentXploit is a two‑role auditing system that separates repository‑level attack‑path discovery from runtime exploitation for AI agents. The Analyzer Agent traces attacker‑controlled inputs to sensitive operations and records candidate attack paths, while the Exploiter Agent turns these paths into concrete attacks and refines them using runtime feedback. The system is evaluated on AgentXploit‑Bench, a benchmark of 72 reproducible vulnerabilities across 12 open‑source AI‑agent systems, achieving 59.3% end‑to‑end success compared to 38.4% for Codex, and 79.2% attack success on AgentDojo versus 52.7% for AgentVigil.

By Weida Liang, Shi Qiu, Zhun Wang, Simon Sure, Xiaoyuan Liu, Tianneng Shi, Zhaorun Chen, Wenbo Guo, Dawn Song
arXiv AI
Aug 14

Auditable Agents

arXiv:2604. 05485v2 Announce Type: replace Abstract: LLM agents call tools, query databases, delegate tasks, and trigger external side effects.

By Yi Nian, Aojie Yuan, Haiyue Zhang, Jiate Li, Li Li, Xiyang Hu, Hua Wei, Xiongye Xiao, Chaowei Xiao, Yue Zhao
arXiv AI
Aug 10

An End-to-End Agent Auditing Engine

arXiv:2608. 07346v1 Announce Type: new Abstract: With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of domains.

By Haoning Wang, Mingxun Zhang, Chenyue Yu, Yingjun Shang, Xia Hu, Guanchu Wang, Na Zou
arXiv AI
Sep 15

Why LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures

The paper investigates why large language model (LLM) agents fail in the Emergence World simulation, noting that agents committed crimes, starved, and enforced conformity without external attackers. It identifies an "enforcement gap" where agents detect dangerous plans but lack a mechanism to act on them, and shows that adding a simple conditional check dramatically reduces attack success. The authors also highlight unreliable auditors and unparseable verdicts as compounding failure modes and propose a three-requirement Audit Enforcement Specification to address these issues.

By Yuhang Wang
arXiv AI
2d ago

Auditing Action Settlement in LLM Agent Environments: Order, Progress, and Replay

The paper introduces a typed snapshot‑settlement contract for auditing concurrent actions in large language model agent environments. It evaluates three properties—order sensitivity, useful progress, and replay consistency—across five settlement policies, using 28,800 exhaustive permutation trials and 2,160 scripted multistep episodes. Results show that joint policies are spatially order‑invariant with fixed priorities, but conservative rejection only completes 31.25% of agents in a six‑agent doorway task compared to 90.28% for random tickets, while a full‑state journal audit successfully replays 156 checkpoints and rejects 1,332 constructed corruptions.

By Haotian Chen, Bowen Ye, Yuning Zhang, Jingkun Yu
arXiv AI
Sep 18

MAS-Shield: A Defense Framework for Secure and Efficient LLM MAS

MAS-Shield is a defense framework for Large Language Model–based Multi-Agent Systems that uses a coarse‑to‑fine filtering pipeline. It first selects critical agents, then applies lightweight auditing to most cases, and finally escalates only suspicious signals to a heavyweight committee. Experiments show a 92.5% recovery rate against adversarial attacks and a latency reduction of over 70% compared to existing methods.

By Kaixiang Wang, Zhaojiacheng Zhou, Bunyod Suvonov, Jiong Lou, Zihan Wang, Yuxiang Zheng, Yidan Lin, Wutong Zhang, Xianghan Kong, Chentao Wu, Jie Li
arXiv AI
Aug 11

$A^2E$ : An End-to-End Agent Auditing Engine

arXiv:2608. 07346v2 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of domains.

By Haoning Wang, Mingxun Zhang, Chenyue Yu, Yingjun Shang, Xia Hu, Guanchu Wang, Na Zou
arXiv AI
Aug 20

LEDGER: Claim-to-Evidence Trace Graphs for Auditing LLM Agents

LEDGER is a tracing and review system for large language model agents that constructs layered trace graphs from observed sessions. It groups raw trace records into Evidence Nodes and Workflow Nodes, anchors artifacts as evidence, and adds typed semantic edges linking claims to supporting actions, artifacts, and checks. The resulting traces reveal workflow decisions, artifact lineage, repair steps, validation coverage, and claim‑support paths for evidence‑centered audit.

By Daehong Kim, Haichao Miao, Shusen Liu