arXiv Machine Learning By Yan Liu, Jie Fu, Tsung-Yi Ho

Safe Evolution with Circuit Anchors

Read the original on arXiv Machine Learning →

arXiv:2608. 05158v1 Announce Type: cross Abstract: In biological evolution, unconstrained mutation can lead to catastrophic outcomes: organisms may evolve enhanced capabilities while losing essential functions for survival.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Sep 2

SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment

SafeEvolve is an experience-driven framework that co‑evolves a harness and policy to improve safety alignment for LLM‑based agents. It uses on‑policy trajectory safety evidence to update safety prompts and hierarchical skills, producing auditable harness artifacts. The policy is trained via a two‑stage SFT‑RL pipeline that bootstraps with the evolved harness and then refines behavior through verifier‑decomposed rewards, yielding a better safety‑utility tradeoff on benchmarks such as AgentDojo.

arXiv AI
Aug 11

SHE: Trajectory-driven Safety Harness Evolution for LLM Agents

arXiv:2608. 09885v1 Announce Type: new Abstract: The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control.

By Wanying Qu, Qinghua Mao, Yu Li, Jiyao Liu, Xin Zhang, Dadi Guo, Yanxu Zhu, Qingyu Liu, Leitao Yuan, Xi Lin, Shanfeng Zhu, Yanwei Fu, Jing Shao, Xia Hu, Dongrui Liu
arXiv AI
6d ago

Evolutionary Safety of Recursive Self-Improving AI: Taxonomy, Risk Discovery, and Evaluation

The paper introduces the concept of Evolutionary Safety for recursive self-improving AI, focusing on how safety properties evolve as an AI system and its successors change. It identifies key risks such as intent drift, error accumulation, and safety-property erosion, and presents a taxonomy covering agent state, model state, evaluation, environment, and update mechanisms. The authors propose methods for discovering and evaluating evolutionary risks, and outline governance principles for modification, selection, authorization, provenance, and recovery, while highlighting open problems for maintaining safety in persistent, adaptive, and recursively self-improving systems.

By Chang Gong, Jingping Bi, Di Yao, Xinjian Liang, Chao Xiang, Ruijie Guo
arXiv AI
2d ago

Multi-Behavioral Evolved Substrates Through Neuromodulation and Activation Selection

The paper investigates whether artificial evolution can replicate biological neuromodulation and diverse neuron types in indirectly encoded substrates. Experiments show that neuromodulation alone cannot overcome a 75% performance ceiling on parity tasks, but combining neuromodulation with per‑task activation function selection allows a single evolving genotype to achieve 100% success across five tasks. This demonstrates that both neuromodulation and evolvable computational primitives are necessary for multi‑behavioral open‑ended evolution.

By Romain Claret, Michael O'Neill, Paul Cotofrei, Kilian Stoffel
arXiv AI
Sep 3

SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment

SafeEvolve is an experience-driven framework that co‑evolves a harness and policy to align large‑language‑model agents with safety goals. It uses completed on‑policy trajectories to update safety prompts and hierarchical skills, then applies a two‑stage SFT‑RL training loop that bootstraps the policy with the evolved harness and refines it through verifier‑augmented rewards. Experiments on agentic safety benchmarks show that SafeEvolve improves the safety‑utility tradeoff, achieving a three‑fold reduction in ASR on AgentDojo for Qwen3.5‑4B while increasing benign utility from 59.79% to 61.86%.

By Qinghua Mao, Wanying Qu, Dadi Guo, Leitao Yuan, Qingyu Liu, Yu Li, Guanxu Chen, Yanwei Fu, Xi Lin, Xia Hu, Dongrui Liu
arXiv AI
Aug 5

A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models

arXiv:2608. 02684v1 Announce Type: cross Abstract: Large Language Models (LLMs) are accelerating biological research, yet this same capability poses a critical biosecurity threat: models that assist in protein engineering can equally be prompted to generate predicted toxin-like sequences, potentially lowering the barrier to biological misuse.

By Shu Quan, Tianfang Hao, Sitong Fang, He Geng, Jiayi Zhou, Boyuan Chen, Kaile Wang, Donghai Hong, Juntao Dai, Yaodong Yang, Jiaming Ji