arXiv AI

SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent

arXiv:2608. 07449v1 Announce Type: new Abstract: LLM agents increasingly adapt to recurring tasks by accumulating procedural knowledge in skills.

arXiv AI
2d ago

Rep2Skill: Representation-Guided Skill Self-Evolution for LLM Agents

Rep2Skill introduces a representation-guided framework that enables large language model agents to self-evolve their textual skills by analyzing internal representation trajectories from agent rollouts. The method identifies execution turns that deviate from successful dynamics and uses these signals, together with execution contexts, as actionable feedback for targeted skill revision. Experiments with two open-source LLMs across two agent environments demonstrate that Rep2Skill consistently outperforms purely text-based approaches, showing that incorporating internal representations can enhance agent self-improvement.

By Kaixing Zhang, Changming Li, Yingdong Shi, Zheng Zhang, Kaitao Song, Wenjie Shi, Jingang Wang, Kan Ren
arXiv Machine Learning
22h ago

SkillSpec: Consensus-Gated Agent Skill Evolution via Representation Specialization

SkillSpec is a two‑phase framework for evolving natural‑language skills in large language model agents. The first phase, consensus‑gated evolution, generates candidate skills from complementary editing intents and commits updates only when paired evaluations reach consensus on overall improvement and non‑negative aggregate gain. The second phase, representation specialization, uses signals from the optimization trajectory to choose an appropriate flat, graph, or hybrid structure for the skill, improving success rates by an average of 6.89% over SkillOpt across six benchmarks and three target language models.

By Huancheng Chen, Xiaodi Sun, Zhaoqiong Huang, Shenyang Huang Shreya Singhal, Jingwen Lu
arXiv AI
2d ago

Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training

The Agent Error Dataset (AED) presents 50,228 error–diagnosis pairs collected from 9,961 source tasks across 33 environments, 19 harness families, and 23 policy models in text‑based agent systems. A five‑stage Agentic Error‑to‑Training (AET) pipeline generates diagnoses and proposed corrections, verifies them against recorded evidence, and creates separate training views for diagnosis and actor recovery. Experiments show that first‑proposal corrections improve verifier pass rates from 18.4% to 51.1%, and fine‑tuning with full‑diagnosis data raises Qwen3‑8B’s exact‑step agreement from 47.2% to 63.6% on a holdout set.

By Kunlun Zhu, Xuyan Ye, Yibo Li, Cheng Qian, Beibin Li, Heng Ji
arXiv AI
5d ago

SkillEvoReg: Regularizing Agent Skill Evolution Against Overfitting

SkillEvoReg is a regularization framework designed to mitigate overfitting in language-model agents that evolve reusable external skills. It combines training-time skill dropout, complexity-aware local regularization, and causal counterexample validation to control skill-state growth and detect regressions. Applied across SkillOpt, SkillEvolBench, and ContinualSkillBench, it preserves downstream performance while improving transfer and later-stage evolution outcomes.

By Guanyu Nie, Fangzhou Zhu, Shixiong Kai, Xiongwei Han, Tao Zhong, Mingxuan Yuan