arXiv:2606. 09316v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) enables agents to access external knowledge at inference time, but it primarily retrieves fragmented declarative evidence, leaving agents to repeatedly infer task procedures from passages, manuals, examples, logs, or trajectories.
By Qianjun Pan, Yutao Yang, Junsong Li, Jie Zhou, Kai Chen, Xin Li, Qin Chen, Liang He
Corpus2Skill is a retrieval architecture that transforms an enterprise knowledge base into a hierarchical skill directory, enabling an LLM agent to navigate from high-level summaries to specific documents and backtrack when necessary. On an enterprise customer‑support benchmark, it outperforms single‑shot dense, hybrid, hierarchical‑retrieval, and agentic RAG baselines in answer quality and grounding, with a moderate cost tradeoff. An eleven‑dataset study shows that corpus navigation excels on single‑domain corpora with a recoverable topical taxonomy but is less effective on open‑domain factoid pools or homogeneous‑tabular corpora, providing a design guideline for knowledge‑grounded systems.
By Yiqun Sun, Pengfei Wei, Lawrence B. Hsieh
mimeo is an open‑source tool that compiles a public expert’s work into a file an agent can load, verifying each quotation against the source text. In experiments, mimeo enabled agents to answer all 20 obscure, quotation‑heavy questions and avoided misstatements that occurred with personas generated from model memory. While it improved knowledge access and reduced misidentification, it did not demonstrate clear transfer of expert judgment or outperform simple keyword search in all metrics.
By Timothy Kassis
SkillSpec is a two‑phase framework for evolving natural‑language skills in large language model agents. The first phase, consensus‑gated evolution, generates candidate skills from complementary editing intents and commits updates only when paired evaluations reach consensus on overall improvement and non‑negative aggregate gain. The second phase, representation specialization, uses signals from the optimization trajectory to choose an appropriate flat, graph, or hybrid structure for the skill, improving success rates by an average of 6.89% over SkillOpt across six benchmarks and three target language models.
By Huancheng Chen, Xiaodi Sun, Zhaoqiong Huang, Shenyang Huang Shreya Singhal, Jingwen Lu
arXiv:2606. 29538v1 Announce Type: cross Abstract: Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge.
By Yijia Fan, Zonglin Di, Zimo Wen, Yifan Yang, Mingxi Cheng, Qi Dai, Bei Liu, Kai Qiu, Yue Dong, Ji Li, Chong Luo
arXiv:2606. 04781v1 Announce Type: new Abstract: Agent Skills today consist largely of free-form prose requiring the agent to read, interpret, and re-derive how to act in every session.
By Zachary Blumenfeld, Jim Webber
arXiv:2607. 02432v1 Announce Type: new Abstract: Scalable and reliable grading of command-line examinations remains a challenge in computing education, where rising enrolments make manual marking difficult and rule-based autograders cannot handle partial credit, equivalent solutions, or syntactic variation.
By Manuel Alonso-Carracedo, Ruben Fernandez-Boullon, Pedro Celard, Francisco J. Rodriguez-Martinez, Lorena Otero-Cerdeira
The paper introduces a structured approach to extracting skill and responsibility level pairs from free text using the Skills Framework for the Information Age (SFIA). It evaluates five methods—including lexical baselines, retrieval‑augmented generation, and multi‑agent crews—against expert‑mapped European ICT role profiles, finding that generative strategies are more precise and that only explicit level‑prediction strategies reliably assign responsibility levels. The study also releases an automated SFIA‑9 corpus and establishes the first reproducible baseline for level‑aware skill extraction.
By Ranuga Disansa, U. S. Samarasinghe, Lasith Gunawardena
arXiv:2609.39325v1 Announce Type: new
Abstract: The ability of Large Language Model (LLM) agents to complete daily and professional work is receiving increasing attention. Training such agents requir...
By Xinyu Zhu, Fenyi Liu, Yuzhu Cai, Shuo Tang, Rui Ye, Linfeng Zhang, Siheng Chen
arXiv:2608. 03451v1 Announce Type: new Abstract: Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multimedia.
By Boyan Li, Zhuowen Liang, Yupeng Xie, Xiaotian Lin, Tianqi Luo, Xinyu Liu, Yizhang Zhu, Zhangyang Peng, Yuan Li, Zhengxuan Zhang, Jiayi Zhang, Nan Tang, Guoliang Li, Yuyu Luo
arXiv:2609.33153v2 Announce Type: replace-cross
Abstract: Reported improvements from tools and reusable skills in large language model agents refer to different comparisons. This critical narrative r...
By Shuyang Zhang
arXiv:2606. 06546v1 Announce Type: new Abstract: Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know.
By Tao Liu, Ye Lu, Ruohua Zhang, Siyu Song, Wentao Liu, Aimin Zhou, Hao Hao