arXiv:2606. 08049v1 Announce Type: new Abstract: AI agents increasingly turn past experience into reusable artifacts such as code, workflows, and procedural memories.
By Amine El Hattami, Nicolas Chapados, Christopher Pal
SkillGLoW introduces a new way for large language model agents to self‑improve by consolidating procedural skills shared across related tasks. Instead of storing all skills in a single global document or a flat per‑task pool, SkillGLoW aggregates local skills into procedural families, compresses them into de‑instantiated global priors, and regenerates instance‑specific details on demand. Experiments on four diverse benchmarks show that these priors improve performance by an average of 17.2 points over a no‑skill baseline, are more compact than per‑task pools, and enable better transfer to unseen tasks.
By Ao Yan, Xin Zhang, Jiawei Du, Joey Tianyi Zhou
SkillGym is an automatic pipeline that generates verifiable environments for training skill-use agents. It crawls internet skills, filters for reproducible workflows, and uses a builder‑reviewer process to create difficulty‑controlled tasks with reference solutions and verifiers. The system builds 6.8k environments, collects 19k successful trajectories, and fine‑tunes LLMs from 2B to 122B parameters, improving performance and skill invocation rates.
By Renxi Wang, Mingshan Hee, Fajri Koto, Timothy Baldwin, Haonan Li
arXiv:2609.27717v1 Announce Type: new
Abstract: Human-written agent skills encode rich workflows for real-world problem solving, but are typically used as external inference-time instructions rather...
By Zhilong Ge, Yuting Shao, Yutao Yang, Yuxuan Cai, Jie Zhou, Kai Chen, Bo Zhang, Qin Chen, Liang He
arXiv:2606. 01311v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly rely on reusable external skills to solve long-horizon interactive tasks.
By Zhuoyun Yu, Xin Xie, Wuguannan Yao, Chenxi Wang, Lei Liang, Xiang Qi, Shumin Deng
arXiv:2607. 23123v1 Announce Type: new Abstract: Existing evaluations of large language models cover knowledge, reasoning, coding, and tool use, but they rarely treat a verifiable deliverable produced within a constrained workflow as the unit of evaluation.
By Summer Sun (Shaqiu Community)