arXiv Computation and Language

Skill-as-Pseudocode: Refactoring Skill Libraries to Pseudocode for LLM Agents

The paper introduces Skill-as-Pseudocode (SaP), a method that automatically converts markdown skill libraries for large language model agents into typed pseudocode with deterministic quality control. SaP extracts typed contracts from clusters of procedural passages and verifies them with a four‑check verifier before inlining them into a rewritten skill skeleton that includes both a typed signature and a concrete action template. On the ALFWorld unseen split, SaP outperforms the Graph-of-Skills baseline, achieving 82/402 paired game wins versus 47/402, while reducing input tokens and LLM calls per game.

arXiv AI
Jun 3

SkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at Scale

arXiv:2606. 03056v1 Announce Type: new Abstract: As LLM agents adopt large skill libraries, selecting the right subset becomes a structural problem rather than a similarity-matching one: skills depend on, conflict with, specialize, or duplicate one another, a structure invisible to both full enumeration and embedding similarity.

By Tong Bai, Zhenglin Wan, Pengfei Zhou, Xingrui Yu, Wangbo Zhao, Yang You, Ivor W. Tsang
arXiv AI
Aug 20

SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents

SkillGate is a method that trains agents to select the correct skill from a large slate during an episode by separating credit signals for skill selection and execution. It addresses the problem of selector credit starvation, where traditional outcome-rewarded RL fails to give sufficient credit to the skill-naming tokens, especially in long-horizon tasks. Experiments on five benchmarks show that SkillGate improves a 9B policy’s success rate from 40.8% to 53.2%, reduces exposure to misleading candidates, and requires fewer skill reads.

By Qingyao Li, Wenxiang Jiao, Shuai Shao, Kangning Zhang, Yuan Lu, Yi Guo, Weiwen Liu, Weinan Zhang, Yong Yu
arXiv AI
Aug 12

CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification

arXiv:2604. 01687v3 Announce Type: replace Abstract: Anthropic proposes the concept of skills for LLM agents to tackle multi-step professional tasks that simple tool invocations cannot address.

By Hanrong Zhang (Steve), Shicheng Fan (Steve), Henry Peng Zou (Steve), Yankai Chen (Steve), Zhenting Wang (Steve), Jiayu Zhou (Steve), Chengze Li (Steve), Wei-Chieh Huang (Steve), Yifei Yao (Steve), Kening Zheng (Steve), Xue (Steve), Liu, Xiaoxiao Li, Philip S. Yu
arXiv AI
Jun 9

Anything2Skill: Compiling External Knowledge into Reusable Skills for Agents

arXiv:2606. 09316v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) enables agents to access external knowledge at inference time, but it primarily retrieves fragmented declarative evidence, leaving agents to repeatedly infer task procedures from passages, manuals, examples, logs, or trajectories.

By Qianjun Pan, Yutao Yang, Junsong Li, Jie Zhou, Kai Chen, Xin Li, Qin Chen, Liang He
arXiv AI
Sep 21

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design

Designer‑RSI presents a continual adaptation framework that lets a frozen frontier model operate professional design software while an external procedural memory learns natural‑language design skills from user traffic. Over five rounds on 1,406 real briefs and 1,869 graded trajectories, the memory grew from 76 to 139 skills, boosting execution success from 72.7% to 99.3% and improving win rates on four design benchmarks. The study shows that widening and deepening the memory, especially together, significantly outperforms a no‑skill baseline.

By Hongyang Du, Lan Yan, Christian Flores, Asim Kadav