arXiv:2609.24663v1 Announce Type: new
Abstract: Self-evolving agents convert interaction feedback into persistent artifacts, such as memories or skills, which in turn guide subsequent decisions. As t...
By Hongqiang Lin, Chao Liu, Xiaofan Bai, Xuan Jin, Yuhong Li, Nenggan Zheng, Xipeng Cao
EVOHARNESSBENCH is a new benchmark that tests how LLM-based agents handle changes in their tool, skill, and agent harnesses over time. It includes 17 deterministic harness streams with 802 tasks, 520 tools, 42 skills, and 62 agents, and evaluates agents in two settings: deployment evaluation and self‑evolving adaptation evaluation. The study finds that harness expansion can cause forgetting, adaptation gains are inconsistent, and preserving old competence does not always aid new capability adaptation, highlighting harness evolution as a distinct challenge for agent development.
By Zixuan Ke, Vaidehi Patil, Haizhou Shi, Yang Li, Ye Liu, Sarath Shekkizhar, Anurag Koul, Jiayu Wang, Xuan Phi Nguyen, Semih Yavuz, Mohit Bansal, Shafiq Joty
arXiv:2609.04280v2 Announce Type: replace-cross
Abstract: Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they observe and what the...
By Zixuan Ke, Vaidehi Patil, Haizhou Shi, Yang Li, Ye Liu, Sarath Shekkizhar, Anurag Koul, Jiayu Wang, Xuan Phi Nguyen, Semih Yavuz, Mohit Bansal, Shafiq Joty
Agent self-evolution in long-horizon LLM systems is largely procedural: useful experience is not merely stored information, but reusable procedures for searching, debugging, and verification. Yet current evaluations do not isolate this form of transfer.
Many important forms of human learning begin with a vague goal, such as "become a better physicist" or "improve at research." Learners must interpret the goal, identify capability gaps, decide how to...
arXiv:2607. 05202v1 Announce Type: new Abstract: Agent self-evolution in long-horizon LLM systems is largely procedural: useful experience is not merely stored information, but reusable procedures for searching, debugging, and verification.
By Xingze Gao, Chuanrui Hu, Hongda Chen, Pengfei Yao, Zhao Wang, Yi Bai, Zhengwei Wu, Yunyun Han, Xiaofeng Cong, Jie Gui, Yafeng Deng, Teng Li
The paper introduces ASPIRE, a benchmark that challenges language model agents to self‑evolve from vague, natural‑language goals without explicit evaluation metrics. In ASPIRE, agents must interpret the goal, select data and update strategies, and decide when to evaluate, all while the downstream tasks remain hidden. Experiments show that while agents can complete training loops, weight‑level improvements are sparse and unstable, and the best evolved harness still falls short of a strong engineered baseline.
By Yuhao Wu, Jingyuan Zhang, Jiajun Shi, Yuxuan Zhang, Xinping Lei, Junting Zhou, Zexuan Wang, Yuchen Wu, Huan Zhou, Duo Wang, Yinzhu Piao, Yongchang Peng, Yunfeng Shi, Jin Chen, Zuo Wang, Jinkai Liu, Jiaheng Liu, Wenxuan Zhang, Shen Yan, Wenhao Huang, Ge Zhang
arXiv:2606. 01314v1 Announce Type: new Abstract: Recent self-evolving agents have shown that skills can be discovered, refined, and accumulated through execution.
By Yangbo Wei, Zhen Huang, Shaoqiang Lu, Junhong Qian, Qifan Wang, Chen Wu, Lei He
arXiv:2608. 16544v1 Announce Type: cross Abstract: Agents increasingly rely on reusable skills to encode task knowledge, tool-use procedures, and validation rules.
By Jianming Chen, Xuanbin Ye, Yawen Wang, Junjie Wang, Qing Wang, Fanjiang XU
arXiv:2609.35596v2 Announce Type: replace-cross
Abstract: Self-evolving LLM agents have gained prominence for their ability to improve after deployment by modifying their harness, including their con...
By Saswat Das, Parvati Viswanathan, Daniel Donnelly, Chang Huang, Sahar Abdelnabi, Ferdinando Fioretto
SEABench is a benchmark designed to study endogenous misalignment in self‑evolving large language model agents. It contains 48 longitudinal task sequences across various evolution surfaces, task domains, and harm types, and includes an adaptive trajectory discovery pipeline that probes for failures while preserving task intent. Evaluations show that self‑evolution improves task completion rates but often introduces safety failures absent in non‑evolving baselines, with divergent safety behaviors reflected in agents’ chain‑of‑thought reasoning that can be monitored to mitigate unsafe actions.
arXiv:2608. 02636v1 Announce Type: cross Abstract: Self-evolving skill systems promise to improve agents by turning execution feedback into persistent skill updates without changing the underlying model.
By Yuxuan Liu, Zhaochen Su, Yuhao Zhang, Jiahe Guo, Zhongwei Xie, Huihao Jing, Lingyun Xie, Qing Zong, Yauwai Yim, Zhixiong Zhang, Haoran Li, Yangqiu Song