arXiv AI By Charles L. Wang, Keir Dorchen, Peter Jin

Agentic Safety is an Epistemic Property, Not a Behavioral One

Read the original on arXiv AI →

arXiv:2606. 28347v1 Announce Type: cross Abstract: Contemporary AI safety spans pre-training interventions, post-training alignment, deployment-time controls, monitoring, and red-teaming.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

Evolutionary Safety of Recursive Self-Improving AI: Taxonomy, Risk Discovery, and Evaluation

The paper introduces the concept of Evolutionary Safety for recursive self-improving AI, focusing on how safety properties evolve as an AI system and its successors change. It identifies key risks such as intent drift, error accumulation, and safety-property erosion, and presents a taxonomy covering agent state, model state, evaluation, environment, and update mechanisms. The authors propose methods for discovering and evaluating evolutionary risks, and outline governance principles for modification, selection, authorization, provenance, and recovery, while highlighting open problems for maintaining safety in persistent, adaptive, and recursively self-improving systems.

By Chang Gong, Jingping Bi, Di Yao, Xinjian Liang, Chao Xiang, Ruijie Guo
arXiv AI
2d ago

Safety in Self-Evolving Agents: A Survey

The article surveys safety concerns for self‑evolving agents that continually update their internal state, such as model parameters and memories, from new interactions. It introduces the SAVER framework, which tracks reusable influence, adaptation, violations, exposure, and response to assess whether safety properties persist as agents evolve. The survey finds that legitimate state can become unsafe when its persistence, authority, or scope expands beyond its original conditions, and highlights gaps in current research on descendant repair and longitudinal evaluation.

By Jiahao Chen, Zhou Feng, Oubo Ma, Yichen Yan, Ruixiao Lin, Hangtao Zhang, Linkang Du, Hengyu An, Yong Yang, Jun Liu, Junhao Li, Naen Xu, Chunyi Zhou, Yuan Su, Zehao Jin, Qianli Ma, Leyi Qi, Yiming Wang, Zhe Ma, Yuwen Pu, Mengyao Du, Yuanyi Song, Enhao Huang, Zhihui Fu, Jun Wang, Jinfeng Li, Yuefeng Chen, Hui Xue, Yiming Li, Tianyu Du, Shouling Ji
arXiv AI
Aug 26

RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards

RePolicy is a reinforcement learning approach designed to invoke safety policies for language model agents by evaluating entire execution trajectories within context-dependent policy libraries. It generates policy-grounded rationales and safety judgments, and is initialized with the PolicyTraj-20K dataset before fine-tuning via GRPO with verifiable rewards and policy-context perturbation. Experiments on six safety benchmarks demonstrate strong safety-detection performance and robust policy invocation across varying contexts.

By Houcheng Jiang, Boxuan Zhang, Qiyong Zhong, Junfeng Fang, Xiang Wang, Xiangnan He
arXiv AI
Aug 11

SHE: Trajectory-driven Safety Harness Evolution for LLM Agents

arXiv:2608. 09885v1 Announce Type: new Abstract: The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control.

By Wanying Qu, Qinghua Mao, Yu Li, Jiyao Liu, Xin Zhang, Dadi Guo, Yanxu Zhu, Qingyu Liu, Leitao Yuan, Xi Lin, Shanfeng Zhu, Yanwei Fu, Jing Shao, Xia Hu, Dongrui Liu