arXiv:2608. 08640v1 Announce Type: new Abstract: Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge.
By Donghong Jiang, Endian Lin, Luoping Cui, Hanqing Liu, Mingjie Liu, Fan Yang, Hong Wang, Zhao Yang, Chuang Zhu
Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retrieving the appropriate skill from a large- scale library remains challenging because realistic user re- quests are often concise and underspecified, stating only the task goal while leaving the required capabilities and execu- tion steps implicit.
arXiv:2605. 17110v2 Announce Type: replace Abstract: Query clustering organizes queries into groups that reflect shared latent capability demands, enabling capability-aware LLM evaluation.
By Fangzhou Wu, Sandeep Silwal, Qiuyi Zhang
arXiv:2609.35816v1 Announce Type: cross
Abstract: Large language model search agents are often trained with synthetic questions whose difficulty is increased through larger evidence graphs, additiona...
By Linzhi Peng, Hanting Chen, Heng Chang, Ke Cheng, Bowen Du, Weifeng Lv
arXiv:2606. 02109v1 Announce Type: new Abstract: Enterprise AI systems that translate natural language into SQL queries and orchestrate multi-step agentic reasoning pipelines require evaluation approaches fundamentally different from academic benchmarks.
By Shannon Serrao, Soumitra Chatterjee, Dorina Strori, Abhishek Sharma, Nathan Miller
arXiv:2606. 30775v1 Announce Type: cross Abstract: Enterprise AI agents route user queries to specialized skills by matching queries against natural language skill descriptions.
By Yangqiaoyu Zhou, Mohammad Alqudah, Kwei-Herng Lai, Aaron Halfaker, Yingqi Xiong, Yaar Harari
arXiv:2607. 22639v1 Announce Type: new Abstract: Parametric retrieval enables LLMs to retrieve tools implicitly by assigning each API a unique virtual token and training the model to generate it via constrained beam search.
By Sai Shruthi Sistla, Ashutosh Hathidara, Christopher Toukmaji, Mayank Shrivastava, Karthikeyan Asokkumar
Agent Seer is a pipeline that automatically synthesizes realistic evaluation scenarios for AI agents that use external tools, using only the tool’s specification (function names, natural‑language descriptions, and typed parameter schemas). Starting from a single Model Context Protocol (MCP) specification, it enriches raw schemas, generates graded scenarios with synthetic tool outputs, and expands them into mock‑data‑grounded multi‑turn dialogues that demonstrate strong tool‑calling correctness and conversational coherence. Across seven diverse MCP specifications, the pipeline achieves high quality, with parameter‑schema complexity emerging as the main driver of quality variation and argument‑value accuracy identified as the dominant failure mode.
By Harish Karumuri, Mahesh Vemula, David Lopes Pegna
arXiv:2602.00994v3 Announce Type: replace
Abstract: Agentic Reinforcement Learning (ARL) trains large language models to interleave reasoning with external tool execution to solve complex tasks. Most...
By Yu Li, Mingyang Yi, Xiuyu Li, Ju Fan, Fuxin Jiang, Binbin Chen, Peng Li, Jie Song, Tieying Zhang
arXiv:2608. 16502v1 Announce Type: new Abstract: Large-scale agents increasingly rely on retrieval to access external capabilities.
By Yiqi Liu, Joseph James, Yang Wang, Chenghao Xiao, Chenghua Lin
CoHyDE is an iterative co‑training framework that jointly trains a dense encoder and an LLM rewriter for tool retrieval from large API catalogs. The encoder is fine‑tuned with InfoNCE on catalog‑style hypothetical descriptions generated by the rewriter, while the rewriter is preference‑aligned via DPO against the encoder’s retrieval scores. On a 10k‑tool subset of ToolBench, three rounds of CoHyDE outperform the best single‑component baseline by 2.5 pp NDCG@5 on standard queries and 6.3 pp on vague queries, with the largest gains on the hardest vague tier.
By Vaishali Senthil, Ashutosh Hathidara, Sebastian Schreiber
arXiv:2608. 07700v1 Announce Type: new Abstract: Translating a natural-language question into a SPARQL query that can be executed against a large knowledge graph requires resolving lexical ambiguity, grounding surface terms in the target ontology, and producing graph patterns that are both syntactically valid and semantically faithful.
By Tommaso Soru, Abdulsobur Oyewale