arXiv AI

What Can Be Enforced? A Theory of Certified Runtime Safety for Tool-Using Agents

arXiv:2607. 22868v1 Announce Type: new Abstract: Runtime guardrails act before irreversible tool calls, but their guarantees depend on what policy state is representable, what a judge observes, and whether intervention changes future behavior.

arXiv AI
Sep 25

Stale Does Not Mean Unsafe: Guard Precision for Tool-Using LLM Agents under Infrastructure State Races

The paper investigates how tool‑using language‑model agents can safely commit changes to infrastructure when external state may change between read and commit. By distinguishing invalidating races from predicate‑preserving and irrelevant ones, the authors evaluate three commit‑time guard granularities—global epoch, read‑set version, and semantic commit predicate—using a deterministic simulator and three quantized model families. The study finds that only the complete predicate guard consistently eliminates unsafe commits, while freshness‑based guards block a large proportion of benign races and model‑side signals fail to replace precise semantic enforcement.

By Zihao Zheng, Jiayu Long, Baichuan Li, Junyi Yao
arXiv AI
Sep 24

Algorithmic Unverifiability of Safety for Fixed and Recursively Self-Improving Systems

The paper proves that algorithmic safety verification for Turing‑complete, self‑modifying systems—whether fixed or recursively self‑improving—is fundamentally limited. Statistically, no verifier can be sound, complete, and tractable across unbounded domains, all finite configurations, or succinctly described finite environments, due to Rice’s, Gödel’s, Trakhtenbrot’s, coNP, and PSPACE barriers. Dynamically, even a single self‑modification step can render safety properties undecidable, and no total supervisory algorithm can guarantee correctness for all such transformations, though a monitor that raises alarms on violations remains feasible. whyItMatters":"The results show that formal safety guarantees for recursive self‑improvement are unattainable, highlighting intrinsic verification barriers for advanced AI systems."

By Jose Pascual Gumbau Mezquita
arXiv AI
Aug 28

Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents

The paper demonstrates that safety mechanisms for autonomous large language model agents fail to compose across iterative loops, as trajectory‑scoped monitors cannot detect attacks whose evidence is spread over multiple iterations. It introduces LoopHarness, a system that maintains a persistent, non‑decaying safety state across loops, bounding unauthorized actions with a constant that does not grow with the number of iterations. The authors provide a comprehensive evaluation protocol, including attacks that require cross‑iteration evidence, module ablations, and adaptive white‑box red‑team testing.

By Chenhao Wu, Haoxuan Jia, Yang Liu, Yingguang Yang, Yuhan Lin, Chongyang Zhang, Hao Zheng, Yulin Huang, Jianshen Zhang, Yongzhi Qi, Shang Luo, Kefu Xu, Jifeng Zhu, Bin Chong
arXiv AI
Sep 17

Autonomy in Check: Governor-Mediated Adaptive Security at the Edge

The paper proposes a split‑control architecture for adaptive security at the network edge, where an untrusted planner emits typed security intents that are vetted by a deterministic governor before being enacted. The governor checks each intent against safety, resource, temporal‑stability, and proportionality invariants, issuing signed receipts for admitted actions that are compiled into eBPF map updates. Experiments on a Raspberry Pi 5 connected to a university 5G test network show the governor can admit, reject, and bound intents at microsecond cost without disrupting protected‑flow regularity.

By Ijaz Ahmad, Ijaz Ahmad, Flavio Esposito, Erkki Harjula