arXiv AI

Repo2Skill-Evo: Repository Skills Go Stale in Silence

arXiv AI
Sep 3

Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

The paper introduces Repo-To-Skill, a method for converting GitHub repositories into reusable AI skills. By distilling operational knowledge from over 1,000 machine‑learning repositories, the authors build the AREX‑Skill Library with more than 5,000 verified skills across 20 areas. Integrating these skills into a research agent—DisCo—yields significant performance boosts on multiple benchmarks, demonstrating the value of reusable, task‑agnostic knowledge.

By Jianlyu Chen, Yuyang Hu, Hongjin Qian, Jiawei Liu, Wenqing Wei, Xiaolong Chen, Defu Lian, Zhicheng Dou, Chaozhuo Li, Qiwei Ye, Zheng Liu
arXiv AI
Sep 17

Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents

The paper introduces EvoSkill-GUI, a training‑free framework that enables GUI agents to evolve their skills during deployment. Each skill is packaged with metadata, executable plans, and recovery rules, and the system follows a reflect‑revise‑reuse loop where the agent instantly revises skills based on execution feedback. Experiments on MobileWorld, AndroidWorld, and OSWorld show consistent performance gains up to +16.2% without any additional training.

By Bofan Chen, Boxuan Zhang, Fei Tang, Zhengxi Lu, Yong Du, Tongbo Chen, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen
arXiv AI
Jul 24

Workflow-Localized Mechanism Learning: Attribution-Guided Repair and Knowledge Reuse for Structured Agent Skills

arXiv:2607. 20999v1 Announce Type: new Abstract: Agent Skills package reusable procedural knowledge as external artifacts for frozen language-model agents, yet existing optimizers do not jointly resolve where a failure occurs in a workflow, which mechanism caused it, and how relevant knowledge from third-party Skills should be reused locally.

By Zibin Lin, Shengli Zhang, Taotao Wang, Yihan Xia, Deen Ma, Guofu Liao
arXiv AI
Sep 15

K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments

K-Bench is a new benchmark designed to evaluate large language model (LLM) unlearning when the models are deployed as agents. Unlike previous benchmarks that only inspect the final answer, K-Bench examines all six channels of a ReAct agent—including chain-of-thought, tool calls, tool observations, and elicited summaries—to determine if a secret is leaked. The benchmark measures leakage for secrets placed in the model weights, prompt, or retrieval store, and finds that many existing unlearning methods fail to prevent leaks in deployed agents, especially when secrets reside in the prompt or retrieval store.

By Guangsheng Yu, Yanna Jiang, Qin Wang, Baihe Ma, Xu Wang
arXiv AI
Sep 7

From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents

The paper introduces an online skill‑evolution framework that transforms interaction traces and evaluator feedback into a persistent, versioned library of reusable procedures for computer‑use agents. By executing each iteration against a frozen library snapshot, the system updates skills without altering the underlying model parameters. Experiments across four OSWorld domains show that the evolving library consistently outperforms an empty‑library baseline, with gains ranging from 5.7 to 18.6 percentage points, while also revealing domain‑specific temporal stability and challenges in skill retrieval and revision.

By Longtao Hu, Xiao Liang, Linchao Zhu