arXiv:2609. 11030v1 Announce Type: new Abstract: AI agents increasingly act through tools and delegated authority, but general incident repositories rarely capture the mechanisms needed to compare public failures with agent-security evaluations.
By Divyanshu Kumar, Rohith HN, Nitin Aravind Birur, Sahil Agarwal, Prashanth Harshangi
PASTABench introduces a benchmark of 1,139 multi-turn trajectories to evaluate proactive safety monitoring in large language models. It formalizes three dimensions of intervention—whether, when, and what risk—to address gaps in step-level isolation and post-hoc trajectory assessment. The study finds that proactive intervention is largely unsolved, with the best model achieving only 40.74% optimal-timing interventions, and reveals that smaller models’ safety scores are often driven by lexical overfitting rather than true risk comprehension.
By Jiapeng Sun, Yujin Zhou, Han Zhu, Pengcheng Wen, Jiayi Zhou, Sirui Han, Yike Guo
AI agents increasingly act through tools and delegated authority, but general incident repositories rarely capture the mechanisms needed to compare public failures with agent-security evaluations. We...
arXiv:2606. 18356v1 Announce Type: cross Abstract: Tool-using language-model agents introduce security failures that go beyond unsafe text: they can disclose protected objects, write persistent memory, send messages, modify databases, or trigger harmful code and tool effects.
By Yuchuan Tian, Mengyu Zheng, Haocheng Mei, Ye Yuan, Chao Xu, Xinghao Chen, Hanting Chen, Yu Wang
arXiv:2609.24264v1 Announce Type: new
Abstract: Tool-use agent traces identify messages and API calls, but procedural analyses also need explicit units of action and inspectable links to their eviden...
By Songqi Li, Dongqing Li, Zheqiao Cheng
arXiv:2607. 07474v1 Announce Type: cross Abstract: Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not.
By Harry Owiredu-Ashley