Agents often repeatedly solve similar task instances from scratch, leading to unnecessary reasoning cost and long execution traces. Prior work has explored workflow reuse and executable skill induction, but it remains unclear which task scenarios admit procedural skills and how the shared procedural structure should be represented across successful traces.
arXiv:2606. 06893v1 Announce Type: new Abstract: Large language model agents increasingly rely on Skills to encode procedural knowledge, yet high-quality Skills remain costly to hand-write.
By Yuyang Zhang, Xinyuan Han, Xudong Jiang, Run Wang
arXiv:2609.09233v1 Announce Type: cross
Abstract: How can language model agents effectively leverage libraries of reusable knowledge to solve long-horizon tasks? Recent work has increasingly focused...
By Wasu Top Piriyakulkij, Rachel Lawrence, Alicia Curth, Sushrut Karmalkar, Niranjani Prasad
SkillLens introduces a hierarchical skill-evolution framework that organizes skills into a four-layer graph of policies, strategies, procedures, and primitives, allowing retrieval at mixed granularity. The system first retrieves semantically relevant skill seeds, expands them via a degree‑corrected random walk, and uses a verifier to decide whether to accept, decompose, rewrite, or skip each visited unit. This approach enables agents to reuse compatible subskills while locally adapting mismatched components, and theoretical analysis shows sublinear cost under sparse mismatch assumptions, with empirical results on MuLocbench and ALFWorld demonstrating consistent improvements over strong baselines.
By Ziyang Yu, Yongliang Miao, Liang Zhao, Bowen Zhu, Hasibul Haque
arXiv:2607. 25853v1 Announce Type: new Abstract: Skills have become an important abstraction for enabling large language model (LLM) agents to reuse past experience in long-horizon interactive tasks.
By Yu Hao, Jinxuan Cai, Qi Zhang, Yawen Li, Zhiqiang Zhang, Chuan Shi, Cheng Yang
arXiv:2606. 01139v1 Announce Type: new Abstract: Agent skills are procedural artifacts that enable LLM agents to execute workflows, verify constraints, and recover from failures.
By Yuxuan Liu, Zhaochen Su, Lingyun Xie, Yuhao Zhang, Qing Zong, Jiahe Guo, Zhongwei Xie, Yiyan Ji, Yauwai Yim, Hongyu Luo, Xiyu Ren, Ruan Chenyu, Haoran Li, Yangqiu Song
HEXIS is a system that compiles agent skills into extended finite state machines, separating skill knowledge from control flow. It uses local instructions within states to guide reasoning and generation, while explicit transition conditions manage execution progress. The incremental compiler maps skill clauses and tool interfaces to state operations, aligns development traces to identify missing operations, and updates are validated through static checks and replay of traces, resulting in improved success rates and reduced execution tokens across benchmarks.
By Minghao LI
arXiv:2608.23417v1 Announce Type: new
Abstract: Agent skills are reusable procedural artifacts that extend language agents with specialized workflows, tool conventions, and domain behaviors at infere...
By Hengjun Wang, Shuyue Wei, Boyi Liu, Jun Yang, Yongxin Tong
arXiv:2608. 05604v1 Announce Type: cross Abstract: Large Language Models (LLMs) increasingly act as agents whose procedural knowledge is stored in reusable skill packages and loaded at inference time.
By Xingyu Tan, Xiaoyang Wang, Qing Liu, Xiwei Xu, Xin Yuan, Liming Zhu, Wenjie Zhang
The paper introduces the Procedural Graph, a framework that structures procedural knowledge into (procedure, relation, procedure) triplets to guide large language model agents in planning and tool usage. At each decision point, a guidance model uses the local subgraph to bias the agent’s next action, while an LLM refiner self‑evolves the graph by editing its topology based on successful versus failed trajectories. Experiments across datasets and LLMs show that Procedural Graphs consistently outperform memory‑based baselines, and the self‑evolution mechanism further improves performance without manual engineering.
arXiv:2609.33772v2 Announce Type: replace
Abstract: Executable environments are critical for post-training agents on tasks that require tool use and multi-step interaction, but constructing executabl...
By Weiyi Xu, Xiaowen Yang, Wen Da, Hang Xu, Canwei Li, Hongjie You, Pusen Dong, Yucheng Zeng, Zhaokai Luo, Mu Chuan
SkillGym is an automatic pipeline that generates verifiable environments for training skill-use agents. It crawls internet skills, filters for reproducible workflows, and uses a builder‑reviewer process to create difficulty‑controlled tasks with reference solutions and verifiers. The system builds 6.8k environments, collects 19k successful trajectories, and fine‑tunes LLMs from 2B to 122B parameters, improving performance and skill invocation rates.
By Renxi Wang, Mingshan Hee, Fajri Koto, Timothy Baldwin, Haonan Li