arXiv:2609.05677v1 Announce Type: cross
Abstract: Lifelong LLM agents increasingly rely on external skill artifacts as one element for preserving and reusing capabilities over time. These skills (usu...
By Chen Shen, Estevam Hruschka
The paper introduces Repo-To-Skill, a method for converting GitHub repositories into reusable AI skills. By distilling operational knowledge from over 1,000 machine‑learning repositories, the authors build the AREX‑Skill Library with more than 5,000 verified skills across 20 areas. Integrating these skills into a research agent—DisCo—yields significant performance boosts on multiple benchmarks, demonstrating the value of reusable, task‑agnostic knowledge.
By Jianlyu Chen, Yuyang Hu, Hongjin Qian, Jiawei Liu, Wenqing Wei, Xiaolong Chen, Defu Lian, Zhicheng Dou, Chaozhuo Li, Qiwei Ye, Zheng Liu
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills introduces DisCo, a research agent that extracts and verifies operational knowledge from GitHub repositories to create reusable AI skills. The agent produces both task‑agnostic skills—compiled into the AREX‑Skill Library of over 5,000 verified skills from 1,000 repositories—and task‑oriented skills tailored to specific research tasks. When equipped with these skills, the agent achieves significant performance gains across multiple benchmarks, outperforming a skill‑free version by 134.3% on MLE‑bench, 34.4% on PaperBench, 9.2% on FrontierCS, and 14.0% on PassNet.
arXiv:2607. 10113v1 Announce Type: new Abstract: Large language model agents increasingly store reusable procedures outside the model.
By Yubo Li
StagedWorkspace is a versioned workspace designed for knowledge‑work AI agents, ensuring that every parsed view, native file edit, and review diff is explicitly tied to a specific version of the workspace state. By binding parsed records and review diffs to content hashes of native files, the system improves performance on tasks such as OfficeQA and APEX‑Agents, achieving higher pass rates and rubric scores compared to single‑view approaches. The study demonstrates that providing dual parsed/native access and visible diffs enhances agent performance, highlighting workspace state as a key experimental variable for future benchmarks.
By Yining Hua, Hongbin Na, Yifan Zhou, Akshay Kalose, Cyrus Ayubcha, Levi Lian
SpecMine is a large-scale corpus that documents Spec-Driven Development (SDD) artifacts in public GitHub repositories. It includes a broad census of 470,795 spec files from 73,030 repositories linked to 17 tools, a focused census of 98,574 Kiro layout files from 12,910 repositories, and a sweep of 5,992 pull requests across 581 repositories that modify specs. The dataset provides enriched metadata, full commit histories, parsed document structures, and over 2.4 million typed references connecting specs to code, sibling documents, PRs, branches, and issues.
By Shyam Agarwal, Anmol Singhal, Travis Breaux, Bogdan Vasilescu
arXiv:2607. 03780v1 Announce Type: cross Abstract: SkillFab is an agent-native platform for turning missing capabilities into reviewed, reusable Agent Skills.
By Anjie Xu, Yifeng Cai, Yi Li, Zixing Wang, Zhiyu Zhang, Jingfan Chen, Ruohan Xu, Leye Wang
arXiv:2608. 12610v1 Announce Type: new Abstract: There are 56,804 public agent skills today, and teams write many more privately.
By Li Yin (Atlas), Zhi Li (Atlas), Zhan Shi (Atlas), Haoran Zhang (Atlas), Haebin Seong (Atlas), Zhangyang (Atlas), Wang
SkillFlow is an open, multi-stage retrieval system that helps AI agents selectively load relevant skills from a large library of community-contributed SKILL.md definitions. The pipeline uses dense retrieval, two rounds of cross-encoder reranking, and LLM-based selection to balance recall and precision. Evaluations on SkillsBench and Terminal-Bench show that SkillFlow improves performance when high-quality skills are available, but retrieval alone does not help if the corpus lacks executable skills for the target domain.
By Fangzhou Li, Pagkratios Tagkopoulos, Ilias Tagkopoulos
arXiv:2606. 09316v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) enables agents to access external knowledge at inference time, but it primarily retrieves fragmented declarative evidence, leaving agents to repeatedly infer task procedures from passages, manuals, examples, logs, or trajectories.
By Qianjun Pan, Yutao Yang, Junsong Li, Jie Zhou, Kai Chen, Xin Li, Qin Chen, Liang He
The paper introduces Scientific Agent Skills, an open library of 163 procedural knowledge modules across 16 scientific domains such as genomics, cheminformatics, medical imaging, study design, and scientific communication. Each skill is organized as a directory containing a versioned, human‑readable instruction file, reference material, and runnable scripts, allowing a language‑model agent to load the appropriate procedure only when needed. The library is publicly licensed and hosted on GitHub, but the authors report no task‑level evaluation or host selection rate.
By Timothy Kassis, Vinayak Agarwal, Yuhuan He, Darshil Patel, Aubrey M. Brueckner
arXiv:2608. 05204v1 Announce Type: new Abstract: LLM-agent ecosystems are rapidly growing around reusable skills: mixed-modality packages of metadata, natural-language instructions, code, tools, references, and operational workflows.
By Jialuo Chen, Minghe Wang, Lingqi Jiang, Jianan Ma, Xinhao Deng, Xiaohu Du, Ruixiao Lin, Yunhao Feng, Linkang Du, Jingyi Wang