arXiv:2608.21375v1 Announce Type: new
Abstract: Heterogeneous agentic retrieval-augmented generation (RAG) systems increasingly orchestrate external APIs, internal databases, vector stores, and graph...
By Yong-eun Cho
Toollery is a training‑free framework that compresses candidate lists for large language model agents, enabling efficient selection from thousands of skills and tools. It generates user‑intent queries from each skill or tool specification, builds a retrieval index, and limits online selection to a compact top‑k set before the LLM makes its final decision. Evaluations on the SkillRouter benchmark, BFCL‑V4, and a proprietary smart‑cockpit dataset show that Toollery improves recall and end‑to‑end selection while keeping selection costs bounded.
By Xiangxi Tian, Ran Guan
Hybrid Semantic Tool Discovery for Enterprise MCP Gateway presents SCOUT, a system that addresses two major challenges in large language model (LLM) agent tool usage: a context‑engineering bottleneck and a tool discoverability barrier. SCOUT reframes tool exposure as a context‑selection problem, injecting only relevant tools into the model’s context window and providing two MCP meta‑tools—tool_search and execute_tool—to perform hybrid retrieval via BM25 and dense vector search. In production at PayPal, SCOUT cuts MCP tool‑token consumption by 99%, dramatically reducing per‑query inference cost while remaining model‑agnostic and requiring no client‑side changes.
By Olympia Saha, Amy Wang, Srinivasan Manoharan
arXiv:2607. 02116v1 Announce Type: new Abstract: Autonomous AI agents increasingly depend on external knowledge stores, yet most retrieval pipelines provide relevance without durable guarantees of provenance, version identity, integrity, traceability, or point-in-time reconstruction.
By Misha Sulpovar (PromptOwl, LLC), Benn R. Konsynski (Goizueta Business School, Emory University), Qaish Kanchwala (IBM Research), Gabe Goodhart (IBM Research)
arXiv:2608. 12282v1 Announce Type: new Abstract: Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation.
By Ankita Rajaram Naik, Anupama Murthi, Benjamin Elder, Siyu Huo, Raavi Gupta, Abhinav Jain, Praveen Venkateswaran, Abdulhamid Adebayo, Danish Contractor
Cartograph is a federated Model Context Protocol (MCP) proxy that reduces AI agent tool discovery from linear catalog traversal to progressive disclosure, exposing only a few proxy tools instead of all definitions. It uses operator-attested capability cards, a three-layer confusable-cluster analysis called Rift, and a two-stage retrieval process to rank servers before tools. In a 22-server, 374-tool deployment, Cartograph achieves higher recall (R@5 = 0.816 vs. 0.592) and drastically fewer tokens (475 vs. 42,450) for discovery exchanges, with minimal latency overhead.
By Justice Owusu Agyemang, Michael Agyare, Kwame Opuni-Boachie Obour Agyekum, Kwame Agyeman-Prempeh Agyekum, Francisca Adoma Acheampong, Jerry John Kponyo
arXiv:2609.13548v1 Announce Type: new
Abstract: Web agents can utilize reusable tools to reduce the cost and latency of low-level browser interaction, but automatically discovered tool collections ca...
By Xinyun Cao, Adriana Szekeres, Fazle Elahi Faisal
arXiv:2606. 12451v1 Announce Type: new Abstract: Large language models deployed as agents over large tool catalogs face a critical tool-retrieval bottleneck.
By Ashutosh Hathidara, Sai Shruthi Sistla, Sebastian Schreiber, Sahil Bansal
SkillFlow is an open, multi-stage retrieval system that helps AI agents selectively load relevant skills from a large library of community-contributed SKILL.md definitions. The pipeline uses dense retrieval, two rounds of cross-encoder reranking, and LLM-based selection to balance recall and precision. Evaluations on SkillsBench and Terminal-Bench show that SkillFlow improves performance when high-quality skills are available, but retrieval alone does not help if the corpus lacks executable skills for the target domain.
By Fangzhou Li, Pagkratios Tagkopoulos, Ilias Tagkopoulos
The paper introduces Enrich‑Retrieve‑Rank, a scalable method for discovering capabilities in large agent ecosystems. It replaces in‑context routing with an offline enrichment step that converts sparse metadata into searchable profiles, followed by an online retrieve‑then‑rank pipeline that returns a ranked shortlist without invoking candidates. Experiments show that as the number of capabilities grows from 10 to 7,278, the new approach maintains higher top‑1 accuracy and reduces cost by 70× compared to full‑context baselines.
By Nazib Sorathiya, Daniel Zhang, Bardiya Akhbari
SemDHT introduces a certified semantic index for discovering agent-accessible capabilities over exact-key distributed hash tables (DHTs). It uses a two-layer semantic sketch—coarse cells for grouping descriptors and residual codes for refining candidate selection—allowing providers to publish at a bounded set of derived keys while requesters probe precision keys first. Anchor committees certify descriptor-to-key consistency, enabling efficient, consistent discovery with fewer lookups and reduced publication fan‑out, as demonstrated by high recall and significant speedups over locality‑sensitive hashing in real‑world experiments.
By Taotao Wang, Chonghe Zhao, Shengli Zhang, Soung Chang Liew
The paper proposes typed federated artifacts—schema‑validated objects with per‑field privacy and dispute resolution—to enable tool‑routing knowledge sharing among frozen, heterogeneous LLM agents. By replacing flat text prompts with typed fields, the authors achieve near‑centralized routing performance on StableToolBench while reducing data size to 20 MB JSON per client. The study also highlights that a simple TF‑IDF classifier can outperform LLM routing on labeled benchmarks, indicating limitations in current evaluation methods.
By Abhijit Chakraborty, Ni Trieu, Vivek Gupta