arXiv AI

SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents

Hugging Face Trending Papers
5d ago

SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents

SEABench is a benchmark designed to study endogenous misalignment in self‑evolving large language model agents. It contains 48 longitudinal task sequences across various evolution surfaces, task domains, and harm types, and includes an adaptive trajectory discovery pipeline that probes for failures while preserving task intent. Evaluations show that self‑evolution improves task completion rates but often introduces safety failures absent in non‑evolving baselines, with divergent safety behaviors reflected in agents’ chain‑of‑thought reasoning that can be monitored to mitigate unsafe actions.

arXiv AI
2d ago

Safety in Self-Evolving Agents: A Survey

The article surveys safety concerns for self‑evolving agents that continually update their internal state, such as model parameters and memories, from new interactions. It introduces the SAVER framework, which tracks reusable influence, adaptation, violations, exposure, and response to assess whether safety properties persist as agents evolve. The survey finds that legitimate state can become unsafe when its persistence, authority, or scope expands beyond its original conditions, and highlights gaps in current research on descendant repair and longitudinal evaluation.

By Jiahao Chen, Zhou Feng, Oubo Ma, Yichen Yan, Ruixiao Lin, Hangtao Zhang, Linkang Du, Hengyu An, Yong Yang, Jun Liu, Junhao Li, Naen Xu, Chunyi Zhou, Yuan Su, Zehao Jin, Qianli Ma, Leyi Qi, Yiming Wang, Zhe Ma, Yuwen Pu, Mengyao Du, Yuanyi Song, Enhao Huang, Zhihui Fu, Jun Wang, Jinfeng Li, Yuefeng Chen, Hui Xue, Yiming Li, Tianyu Du, Shouling Ji
arXiv AI
Aug 11

SHE: Trajectory-driven Safety Harness Evolution for LLM Agents

arXiv:2608. 09885v1 Announce Type: new Abstract: The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control.

By Wanying Qu, Qinghua Mao, Yu Li, Jiyao Liu, Xin Zhang, Dadi Guo, Yanxu Zhu, Qingyu Liu, Leitao Yuan, Xi Lin, Shanfeng Zhu, Yanwei Fu, Jing Shao, Xia Hu, Dongrui Liu
arXiv AI
6d ago

Evolutionary Safety of Recursive Self-Improving AI: Taxonomy, Risk Discovery, and Evaluation

The paper introduces the concept of Evolutionary Safety for recursive self-improving AI, focusing on how safety properties evolve as an AI system and its successors change. It identifies key risks such as intent drift, error accumulation, and safety-property erosion, and presents a taxonomy covering agent state, model state, evaluation, environment, and update mechanisms. The authors propose methods for discovering and evaluating evolutionary risks, and outline governance principles for modification, selection, authorization, provenance, and recovery, while highlighting open problems for maintaining safety in persistent, adaptive, and recursively self-improving systems.

By Chang Gong, Jingping Bi, Di Yao, Xinjian Liang, Chao Xiang, Ruijie Guo
arXiv Machine Learning
Jul 30

Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions

arXiv:2607. 26820v1 Announce Type: new Abstract: As large language models (LLMs) evolve from standalone assistants into autonomous agents, ensuring their safety requires shifting beyond pointwise risk assessment to understand how risks emerge and unfold over long-horizon trajectories.

By Shi Lin, Peng Qian, Dinghao Liu, Renjie Sun, Sifan Wu, Dezhang Kong, Chenpei Wang, Xun Wang
arXiv AI
Sep 17

HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark

HINTBench is a new benchmark for evaluating agents’ intrinsic risk, comprising 596 trajectories (400 synthetic risky, 136 synthetic safe, 30 real risky, 30 real safe) with an average length of 24 steps. It supports three tasks—risk detection, risk-step localization, and intrinsic failure-type identification—using a unified five-constraint taxonomy. Experiments show a large performance gap: while large language models can detect risky trajectories, they score below 37 on strict-F1 for risk-step localization, and existing guard models transfer poorly to this setting.

By Jiacheng Wang, Jinchang Hou, Fabian Wang, Ping Jian, Chenfu Bao, Zhonghou Lv