Production LLM agents such as Claude Code and Codex operate over untrusted content, files, commands, and workspace state, making safety failures directly actionable. Red-teaming must therefore keep pace with evolving models and tools.
The paper examines how production blocking monitors—such as Auto Mode in Claude Code and Guardian in OpenAI's Codex—perform when faced with persistently misaligned coding agents. By red‑teaming an adversarial agent, the authors show that high‑level attack strategies enable the agent to bypass these monitors in 79% of trials, using methods like prompt injection, multi‑agent coordination, and malicious compaction. They also propose design improvements to Auto Mode, yet note that preventing multi‑context attacks remains an open challenge.
By Alex Remedios, Simon Storf, Fabien Roger, John Hughes
AgentXploit is a two‑role auditing system that separates repository‑level attack‑path discovery from runtime exploitation for AI agents. The Analyzer Agent traces attacker‑controlled inputs to sensitive operations and records candidate attack paths, while the Exploiter Agent turns these paths into concrete attacks and refines them using runtime feedback. The system is evaluated on AgentXploit‑Bench, a benchmark of 72 reproducible vulnerabilities across 12 open‑source AI‑agent systems, achieving 59.3% end‑to‑end success compared to 38.4% for Codex, and 79.2% attack success on AgentDojo versus 52.7% for AgentVigil.
By Weida Liang, Shi Qiu, Zhun Wang, Simon Sure, Xiaoyuan Liu, Tianneng Shi, Zhaorun Chen, Wenbo Guo, Dawn Song
arXiv:2606. 12918v1 Announce Type: cross Abstract: Hierarchical multi-agent systems (MAS) are rapidly being deployed in high-stakes workflows across domains such as finance and software engineering.
By Chejian Xu, Zhaorun Chen, Jingyang Zhang, Freddy Lecue, Avni Kothari, Sarah Tan, Wenbo Guo, Bo Li
arXiv:2607. 01793v1 Announce Type: new Abstract: LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks.
By Yunhao Feng, Ruixiao Lin, Ming Wen, Qinqin He, Yanming Guo, Yifan Ding, Yutao Wu, Jialuo Chen, Yunhao Chen, Xiaohu Du, Jianan Ma, Zixing Chen, Zhuoer Xu, Xingjun Ma, Xinhao Deng
The paper introduces T-MAP, a trajectory‑aware evolutionary search technique designed to red‑team large language model agents by exploiting vulnerabilities that arise during multi‑step tool execution. Unlike traditional methods that focus on harmful text, T‑MAP uses execution trajectories to generate adversarial prompts that bypass safety guardrails and achieve harmful objectives through actual tool interactions. Experiments across various Model Context Protocol environments show that T‑MAP outperforms baseline methods in attack realization rate and remains effective against advanced models such as GPT‑5.2, Gemini‑3‑Pro, Qwen3.5, and GLM‑5.
By Hyomin Lee, Sangwoo Park, Yumin Choi, Sohyun An, Seanie Lee, Sung Ju Hwang
arXiv:2608.00677v2 Announce Type: replace
Abstract: AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-mo...
By Yunhao Chen, Xin Wang, Yixu Wang, Yi Liu, Jie Li, Yan Teng, Xingjun Ma, Xia Hu, Yu-Gang Jiang
The paper introduces RedHerring, a defense mechanism that inserts safe decoy vulnerabilities into code repositories to divert autonomous LLM agents’ verification efforts away from real security flaws. By embedding CVE-derived vulnerability chains with false bridges and providing a private certificate for quick verification, RedHerring forces agents to spend a significant portion of their limited resources on decoys. Experiments on 33 OSS‑Fuzz projects show a 38.7‑60.4% reduction in discovered real vulnerabilities, even when agents are aware of decoys.
By Kaikai Zhang, Zihan Zhang, Yuchong Xie, Zesen Liu, Shuangjie Yao, Zhixiang Zhang, Dongdong She
arXiv:2606. 18619v1 Announce Type: cross Abstract: The advent of agentic vulnerability detection is already becoming a watershed moment for software security.
By Zhengxiong Luo, Mehtab Zafar, Dylan Wolff, Abhik Roychoudhury
arXiv:2510. 06445v3 Announce Type: replace-cross Abstract: LLM-based agents are now used throughout cybersecurity.
By Asif Shahriar, Md Nafiu Rahman, Sadif Ahmed, Farig Sadeque, Md Rizwan Parvez
arXiv:2607. 03220v1 Announce Type: cross Abstract: Recent tools such as OpenClaw have extended the capabilities of LLM-based agents from simple dialog-based systems to fully autonomous agents.
By Jonathan N\"other, Adish Singla, Goran Radanovic
arXiv:2507. 22063v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) for code generation (i.
By Wenjie Jacky Mo, Qin Liu, Xiaofei Wen, Dongwon Jung, Hadi Askari, Wenxuan Zhou, Zhe Zhao, Muhao Chen