arXiv AI

SafeCommit: Certifying When Memory-Grounded Agents May Safely Act

arXiv:2608. 04289v1 Announce Type: new Abstract: Long-horizon agents increasingly use persistent memory and tools to take actions with external side effects.

arXiv AI
Aug 7

DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model

arXiv:2608. 05695v1 Announce Type: new Abstract: As large language model (LLM) agents increasingly invoke external tools and interact with real-world systems, unsafe actions may cause irreversible consequences on external states, user data, and downstream services.

By Wenhao Lin, Chenyu Yu, Xingwei Lin, Sicong Cao, Xiang Chen, Lei Xue, Le Yu, Letian Sha, Chunming Wu
arXiv AI
Sep 25

Stale Does Not Mean Unsafe: Guard Precision for Tool-Using LLM Agents under Infrastructure State Races

The paper investigates how tool‑using language‑model agents can safely commit changes to infrastructure when external state may change between read and commit. By distinguishing invalidating races from predicate‑preserving and irrelevant ones, the authors evaluate three commit‑time guard granularities—global epoch, read‑set version, and semantic commit predicate—using a deterministic simulator and three quantized model families. The study finds that only the complete predicate guard consistently eliminates unsafe commits, while freshness‑based guards block a large proportion of benign races and model‑side signals fail to replace precise semantic enforcement.

By Zihao Zheng, Jiayu Long, Baichuan Li, Junyi Yao
arXiv AI
4d ago

Boundary-State Control for Tool-Using Language-Model Agents: Commit-Time Consistency under State Drift

The paper introduces BSC‑R, a deterministic effect‑boundary mechanism that ties a single‑use commit authorization to the specific action and the semantic state that justified it, aiming to close the proposal‑to‑commit gap in tool‑using language‑model agents. Experiments on 2,847 AgentDojo episodes and 10,302 frozen proposals show that BSC‑R preserves the agent’s original behavior while rejecting unauthorized changes, and further tests on a boundary‑drift experiment and the CONTINUITY suite demonstrate high success rates in valid contexts and robust handling of replay and ambiguous cases. However, broader testing reveals that BSC‑R still allows a 25% invalid‑effect commit rate in a larger attack set, indicating that it provides scoped, not universal, safety.

By Wesley Shu
arXiv AI
Aug 28

Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents

The paper demonstrates that safety mechanisms for autonomous large language model agents fail to compose across iterative loops, as trajectory‑scoped monitors cannot detect attacks whose evidence is spread over multiple iterations. It introduces LoopHarness, a system that maintains a persistent, non‑decaying safety state across loops, bounding unauthorized actions with a constant that does not grow with the number of iterations. The authors provide a comprehensive evaluation protocol, including attacks that require cross‑iteration evidence, module ablations, and adaptive white‑box red‑team testing.

By Chenhao Wu, Haoxuan Jia, Yang Liu, Yingguang Yang, Yuhan Lin, Chongyang Zhang, Hao Zheng, Yulin Huang, Jianshen Zhang, Yongzhi Qi, Shang Luo, Kefu Xu, Jifeng Zhu, Bin Chong
arXiv Computation and Language
Aug 28

The Cold-Start Safety Gap in LLM Agents

The paper investigates whether tool‑calling large language model agents maintain consistent safety throughout a conversation. It finds that agents are most vulnerable at the very start of a session, with safety improving significantly after completing a few regular agentic tasks—a phenomenon termed the cold‑start safety gap. The authors introduce the Safety Over Depth for Agents (SODA) benchmark to systematically study this effect, evaluate multiple models, and demonstrate that warming up agents with regular tasks before deployment enhances safety while preserving utility.

By Chung-En Sun, Linbo Liu, Tsui-Wei Weng
arXiv AI
Jun 6

From Risk Classification to Action Plan Remediation: A Guardrail Feedback Driven Framework for LLM Agents

arXiv:2606. 05805v1 Announce Type: new Abstract: LLM-based guardrails typically safeguard agents by evaluating proposed actions or inputs before execution, producing safety signals such as binary allow/deny decisions, risk categories, and/or explanatory rationales about potential policy violations.

By Yuhao Sun, Jiacheng Zhang, Shaanan Cohney, Zhexin Zhang, Feng Liu, Xingliang Yuan
arXiv AI
Sep 4

Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory

The paper introduces PlanFence, a dependency-scoped action‑validation protocol for distributed large language model (LLM) agent teams. PlanFence requires plans to cite the exact public records they rely on, and executors validate only those records that could affect the pending action, replanning or blocking if validation is incomplete. In 30 controlled live workflows, a freshness‑only executor always acted on obsolete plans, whereas PlanFence completed all tasks without invalid actions, demonstrating controlled safety and system‑cost benefits.

By Evan Chen, Shiqiang Wang, Christopher G. Brinton