arXiv AI By Gloria Felicia (University of Virginia), Zitha Sasindran (Indian Institute of Science Bangalore), Jinfeng He (Cornell University), Michael Eniolade (University of the Cumberlands), Hemant Kumar (University of Arizona), Milan Hussain Angati (California State University Northridge)

StepShield: When, Not Whether to Intervene on Rogue Agents

Read the original on arXiv AI →

arXiv:2601. 22136v2 Announce Type: replace-cross Abstract: Agent safety benchmarks measure whether a monitor detects harm, not when.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 24

PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety

PASTABench introduces a benchmark of 1,139 multi-turn trajectories to evaluate proactive safety monitoring in large language models. It formalizes three dimensions of intervention—whether, when, and what risk—to address gaps in step-level isolation and post-hoc trajectory assessment. The study finds that proactive intervention is largely unsolved, with the best model achieving only 40.74% optimal-timing interventions, and reveals that smaller models’ safety scores are often driven by lexical overfitting rather than true risk comprehension.

By Jiapeng Sun, Yujin Zhou, Han Zhu, Pengcheng Wen, Jiayi Zhou, Sirui Han, Yike Guo
arXiv AI
Jun 18

SafeClawBench: Separating Semantic, Audit-Evidence, and Sandbox Harm in Tool-Using LLM Agents

arXiv:2606. 18356v1 Announce Type: cross Abstract: Tool-using language-model agents introduce security failures that go beyond unsafe text: they can disclose protected objects, write persistent memory, send messages, modify databases, or trigger harmful code and tool effects.

By Yuchuan Tian, Mengyu Zheng, Haocheng Mei, Ye Yuan, Chao Xu, Xinghao Chen, Hanting Chen, Yu Wang