The paper introduces ANCHOR, an external supervisory framework driven by large language models (LLMs) that provides evaluative feedback at multiple stages of self‑evolving agents. By integrating ANCHOR into two open‑source self‑evolving agent frameworks, the authors demonstrate that it significantly improves safety performance while preserving core capabilities across coding, mathematical reasoning, and safety tasks. The study also finds that supervision based on execution results is especially effective and that increasing supervision frequency yields diminishing returns, offering practical guidance for future research.
By Dianxing Shi, Bowen Wang, Junqi He, Junhao Chen, Yuta Nakashima
SEABench is a benchmark designed to study endogenous misalignment in self‑evolving large language model agents. It contains 48 longitudinal task sequences across various evolution surfaces, task domains, and harm types, and includes an adaptive trajectory discovery pipeline that probes for failures while preserving task intent. Evaluations show that self‑evolution improves task completion rates but often introduces safety failures absent in non‑evolving baselines, with divergent safety behaviors reflected in agents’ chain‑of‑thought reasoning that can be monitored to mitigate unsafe actions.
The paper introduces the concept of Evolutionary Safety for recursive self-improving AI, focusing on how safety properties evolve as an AI system and its successors change. It identifies key risks such as intent drift, error accumulation, and safety-property erosion, and presents a taxonomy covering agent state, model state, evaluation, environment, and update mechanisms. The authors propose methods for discovering and evaluating evolutionary risks, and outline governance principles for modification, selection, authorization, provenance, and recovery, while highlighting open problems for maintaining safety in persistent, adaptive, and recursively self-improving systems.
By Chang Gong, Jingping Bi, Di Yao, Xinjian Liang, Chao Xiang, Ruijie Guo
arXiv:2609.35596v2 Announce Type: replace-cross
Abstract: Self-evolving LLM agents have gained prominence for their ability to improve after deployment by modifying their harness, including their con...
By Saswat Das, Parvati Viswanathan, Daniel Donnelly, Chang Huang, Sahar Abdelnabi, Ferdinando Fioretto
arXiv:2606. 28347v1 Announce Type: cross Abstract: Contemporary AI safety spans pre-training interventions, post-training alignment, deployment-time controls, monitoring, and red-teaming.
By Charles L. Wang, Keir Dorchen, Peter Jin
arXiv:2608. 09885v1 Announce Type: new Abstract: The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control.
By Wanying Qu, Qinghua Mao, Yu Li, Jiyao Liu, Xin Zhang, Dadi Guo, Yanxu Zhu, Qingyu Liu, Leitao Yuan, Xi Lin, Shanfeng Zhu, Yanwei Fu, Jing Shao, Xia Hu, Dongrui Liu