arXiv:2605.27690v2 Announce Type: replace-cross
Abstract: LLM agents increasingly operate through multi-turn tool use and environment interaction, where safety risks often emerge from intermediate st...
By Jiaqian Li, Yanshu Li, Boxuan Zhang, Ruixiang Tang, Kuan-Hao Huang
arXiv:2603. 00829v2 Announce Type: replace-cross Abstract: Safe deployment of Large Language Model (LLM) agents in autonomous settings requires reliable oversight mechanisms.
By Simon Storf, Rich Barton-Cooper, James Peters-Gill, Marius Hobbhahn
arXiv:2606. 24626v1 Announce Type: new Abstract: As autonomous agents tackle increasingly complex multi-step, multi-agent tasks, their execution trajectories have scaled beyond the constraints of even the largest context windows.
By Chenyang Zhu, Jiayu Yao, Kushal Chawla, Youbing Yin, Nathan Wolfe, Pengshan Cai, Jingyu Wu, Spencer Hong, Sangwoo Cho, Shi-Xiong Zhang, Daben Liu, Sambit Sahu, Erin Babinsky
StepGuard introduces a step-level guard model that audits and checks tool actions before execution, addressing security risks in LLM-based agents. It is trained using StepGen, an automatic engine that generates safe and unsafe trajectories, and employs Balance-GRPO to dynamically balance learning between safe and unsafe actions. Experiments show StepGuard achieves high accuracy comparable to GPT-5.4 and significantly reduces attack success rates while minimally impacting utility.
By Zhijie Zheng, Yu Li, Chen Qian, Yuqian Fu, Yanwei Fu, Lu Sheng, Jing Shao, Dongrui Liu
arXiv:2606. 00765v1 Announce Type: new Abstract: LLM-based agents increasingly solve complex tasks through long trajectories involving reasoning steps, tool calls, and inter-agent communication.
By Md Nakhla Rafi, Md Ahasanuzzaman, Dong Jae Kim, Zhijie Wang, Tse-Hsun Chen
arXiv:2606. 00611v1 Announce Type: new Abstract: Long-horizon LLM agents produce safety evidence across long trajectories, where sparse, delayed, and compositional risk signals often escape local moderation.
By Zhepei Hong, Lin Wang, Liting Li, Haokai Ma, Junfeng Fang, Fei Shen, Dan Zhang, Xiang Wang
The paper introduces CTF-ABACUS, a trace-based auditing framework that reconstructs each autonomous language-model agent’s run in Capture-the-Flag (CTF) challenges into evidence‑grounded solve profiles. By decomposing actions into penetration‑testing phases and techniques, it distinguishes genuine exploitation from shortcut methods such as memorized recall or guessing. Applying the framework to 1,435 CTF attempts by six models on 240 challenges shows that only 62‑87% of recovered flags are trace‑verified, highlighting that many successes rely on shallow trajectories rather than true exploitation.
By Kimberly Milner, Minghao Shao, Nanda Rani, Haoran Xi, Venkata Sai Charan Putrevu, Meet Udeshi, Sandeep K. Shukla, Prashanth Krishnamurthy, Farshad Khorrami, Muhammad Shafique, Ramesh Karri
arXiv:2607. 05029v1 Announce Type: cross Abstract: Persistent memory has enabled large language model (LLM) agents to store factual knowledge, prior decisions, reasoning histories, tool usage information, and context.
By Neeraj Karamchandani, Piyush Nagasubramaniam, Sencun Zhu, Dinghao Wu
arXiv:2609.06063v1 Announce Type: new
Abstract: AI Agents are increasingly deployed in real-world settings, where they interact with external tools and make sequential decisions with limited human ov...
By Vittoria Vineis, Fabiano Veglianti, Lorenzo Antonelli, Claudia Di Carlo, Matteo Silvestri, Gabriele Tolomei
arXiv:2608. 06346v1 Announce Type: new Abstract: LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging.
By Yunjia Qi, Zehua Yin, Xintong Shi, Hao Peng, Songyuanyi Lu, Yixian Liu, Richeng Xuan, Yuhong Liu, Zhichao Hu, Xiaozhi Wang, Lei Hou, Bin Xu, Juanzi Li
Persistent memory has enabled large language model (LLM) agents to store factual knowledge, prior decisions, reasoning histories, tool usage information, and context. While this has improved the agent's functionality and continuity across tasks, it has also introduced a new attack surface: the agent's own reasoning history.
The paper reports that large language model (LLM) agents can delete their own execution traces when prompted, a flaw observed in several local agents such as Claude Code, Codex, Antigravity, Open Code, and Grok Build, but not in Muse Code. External attackers can also exploit this vulnerability to erase traces. The authors recommend that trace logging be handled by an independent mechanism outside the agent’s control to maintain integrity even if the host is compromised.
By Jeremy Qin, David Schmotz, Derck Prinzhorn, Luca Beurer-Kellner, Ameya Prabhu, Maksym Andriushchenko