arXiv AI

SkillSight: Seeing Through Shared Descriptions for Accurate Skill Retrieval

arXiv:2607. 18785v1 Announce Type: new Abstract: As large language model agents gain access to increasingly large skill libraries, retrieving the right skill becomes critical to reliable capability selection and execution.

arXiv AI
6d ago

SkillFlow: Scalable and Efficient Agent Skill Retrieval System

SkillFlow is an open, multi-stage retrieval system that helps AI agents selectively load relevant skills from a large library of community-contributed SKILL.md definitions. The pipeline uses dense retrieval, two rounds of cross-encoder reranking, and LLM-based selection to balance recall and precision. Evaluations on SkillsBench and Terminal-Bench show that SkillFlow improves performance when high-quality skills are available, but retrieval alone does not help if the corpus lacks executable skills for the target domain.

By Fangzhou Li, Pagkratios Tagkopoulos, Ilias Tagkopoulos
Hugging Face Trending Papers
Aug 9

SkillReason: Reasoning-Enhanced Agent Skill Retrieval for Implicit User Requests

Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retrieving the appropriate skill from a large- scale library remains challenging because realistic user re- quests are often concise and underspecified, stating only the task goal while leaving the required capabilities and execu- tion steps implicit.

arXiv AI
Jun 9

Skill Retrieval Augmentation for Agentic AI

arXiv:2604. 24594v3 Announce Type: replace-cross Abstract: As large language models (LLMs) evolve into agentic problem solvers, they increasingly rely on external, reusable skills to handle tasks beyond their native parametric capabilities.

By Weihang Su, Jianming Long, Qingyao Ai, Qiaozhi He, Yichen Tang, Changyue Wang, Yiteng Tu, Yingbo Wang, Yiqun Liu
arXiv Machine Learning
Sep 2

Field-Aware Agent Skill Retrieval

Field-Aware Agent Skill Retrieval explores how keeping the distinct fields of skill documents—such as name, description, and body—separate can improve retrieval performance. By computing sparse and dense similarities for each field independently and combining them with either uniform weights or a small MLP, the authors achieve higher Recall@10 scores on two benchmarks, SkillRet and SRA-Bench. The study shows that the advantage of field-aware representation grows as the skill bank expands, indicating its importance for large-scale lifelong learning agents.

By Paimon Goulart, Liang Wu, Kelly Wan, Evangelos E. Papalexakis, Liangjie Hong
arXiv Machine Learning
Aug 5

Field Aware Agent Skill Retrieval

arXiv:2608. 02880v1 Announce Type: cross Abstract: As lifelong learning agents accumulate lifelong growing skill banks, retrieving the correct skill becomes an increasingly important bottleneck.

By Paimon Goulart, Liang Wu, Kelly Wan, Evangelos E. Papalexakis, Liangjie Hong
arXiv AI
4d ago

From Lexical Baselines to Agentic Retrieval-Augmented Generation: Structured Skill and Responsibility-Level Extraction with the SFIA Framework

The paper introduces a structured approach to extracting skill and responsibility level pairs from free text using the Skills Framework for the Information Age (SFIA). It evaluates five methods—including lexical baselines, retrieval‑augmented generation, and multi‑agent crews—against expert‑mapped European ICT role profiles, finding that generative strategies are more precise and that only explicit level‑prediction strategies reliably assign responsibility levels. The study also releases an automated SFIA‑9 corpus and establishes the first reproducible baseline for level‑aware skill extraction.

By Ranuga Disansa, U. S. Samarasinghe, Lasith Gunawardena
arXiv Computation and Language
Sep 17

M-SQE: Multilingual Skill Quality Estimation for Enhancing Language Equality in Agentic Skill Use

M‑SQE is a post‑retrieval framework that estimates the quality of multilingual agent skills by combining a Theory view (intrinsic quality) and an Action view (task‑grounded utility) into a domain‑conditioned score. It was evaluated on general, tool‑use, and cultural skill‑use domains, showing a task‑success improvement of at least +3.5 points over baselines across three retrievers. The method notably boosts performance for low‑resource languages, raising Hindi by +12.9 pp and Swahili by +5.6 pp, and achieves strong results across six cultural regions, advancing linguistic and cultural equality in agentic skill use.

By Yilun Liu, Shimin Tao, Minggui He, Chenxin Liu, Li Zhang, Chen Liu, Miao Zhang, Jiaxin Guo, Min Zhang, Liqun Deng, Xiaojun Meng, Daimeng Wei