arXiv AI

OpenSkill: Open-World Self-Evolution for LLM Agents

arXiv:2606. 06741v1 Announce Type: new Abstract: Self-evolving agents requires adaptation after deployment, but existing approaches assume a usable learning loop, such as curated skills, successful trajectories, or verifier signals.

arXiv AI
Aug 12

CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification

arXiv:2604. 01687v3 Announce Type: replace Abstract: Anthropic proposes the concept of skills for LLM agents to tackle multi-step professional tasks that simple tool invocations cannot address.

By Hanrong Zhang (Steve), Shicheng Fan (Steve), Henry Peng Zou (Steve), Yankai Chen (Steve), Zhenting Wang (Steve), Jiayu Zhou (Steve), Chengze Li (Steve), Wei-Chieh Huang (Steve), Yifei Yao (Steve), Kening Zheng (Steve), Xue (Steve), Liu, Xiaoxiao Li, Philip S. Yu
arXiv AI
4d ago

SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation

SkillGym is an automatic pipeline that generates verifiable environments for training skill-use agents. It crawls internet skills, filters for reproducible workflows, and uses a builder‑reviewer process to create difficulty‑controlled tasks with reference solutions and verifiers. The system builds 6.8k environments, collects 19k successful trajectories, and fine‑tunes LLMs from 2B to 122B parameters, improving performance and skill invocation rates.

By Renxi Wang, Mingshan Hee, Fajri Koto, Timothy Baldwin, Haonan Li
arXiv AI
Jul 7

EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer

arXiv:2607. 05202v1 Announce Type: new Abstract: Agent self-evolution in long-horizon LLM systems is largely procedural: useful experience is not merely stored information, but reusable procedures for searching, debugging, and verification.

By Xingze Gao, Chuanrui Hu, Hongda Chen, Pengfei Yao, Zhao Wang, Yi Bai, Zhengwei Wu, Yunyun Han, Xiaofeng Cong, Jie Gui, Yafeng Deng, Teng Li
arXiv AI
3d ago

Rep2Skill: Representation-Guided Skill Self-Evolution for LLM Agents

Rep2Skill introduces a representation-guided framework that enables large language model agents to self-evolve their textual skills by analyzing internal representation trajectories from agent rollouts. The method identifies execution turns that deviate from successful dynamics and uses these signals, together with execution contexts, as actionable feedback for targeted skill revision. Experiments with two open-source LLMs across two agent environments demonstrate that Rep2Skill consistently outperforms purely text-based approaches, showing that incorporating internal representations can enhance agent self-improvement.

By Kaixing Zhang, Changming Li, Yingdong Shi, Zheng Zhang, Kaitao Song, Wenjie Shi, Jingang Wang, Kan Ren
arXiv AI
Sep 16

ANCHOR: An External LLM-Driven Supervisory Module Facilitating Healthy Evolution in Self-Evolving Systems

The paper introduces ANCHOR, an external supervisory framework driven by large language models (LLMs) that provides evaluative feedback at multiple stages of self‑evolving agents. By integrating ANCHOR into two open‑source self‑evolving agent frameworks, the authors demonstrate that it significantly improves safety performance while preserving core capabilities across coding, mathematical reasoning, and safety tasks. The study also finds that supervision based on execution results is especially effective and that increasing supervision frequency yields diminishing returns, offering practical guidance for future research.

By Dianxing Shi, Bowen Wang, Junqi He, Junhao Chen, Yuta Nakashima
Hugging Face Trending Papers
Aug 11

SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models

Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within agent harness systems as an essential mechanism to continually constrain a language model's behavior space for repeatable, high-quality task execution. However, because strong closed-source models entail high inference costs, current popular agent harnesses, such as Codex and OpenClaw, remain prohibitively expensive when deploying these skills to accomplish real-world tasks.