arXiv Machine Learning

Toolcompass: Guiding Tool Trialing, Not Suppressing It

ToolCompass is a post‑training framework that improves how large language model agents explore new tools by organizing tool‑call representations according to shared functions. It models each function class as a von Mises–Fisher distribution, reducing variation within a function while increasing separation between different functions, thereby guiding exploration toward functionally similar unseen tools. Experiments on AppWorld and FTRL show consistent gains, with up to a 10.71‑percentage‑point improvement in out‑of‑distribution task success over vanilla post‑training and outperforming competitive baselines.

arXiv AI
Sep 3

UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents

UniToolCall introduces a unified framework for tool-use in large language model agents, standardizing toolset construction, dataset generation, and evaluation. The framework aggregates over 22,000 tools and creates a hybrid training corpus of more than 390,000 instances by combining ten public datasets with synthetically generated, structurally controlled trajectories. It models diverse interaction patterns—single‑hop vs. multi‑hop, single‑turn vs. multi‑turn, serial vs. parallel execution—and adds an Anchor Linkage mechanism to enforce cross‑turn dependencies, while converting seven public benchmarks into a common Query–Action–Observation–Answer format for fine‑grained evaluation.

By Yijuan Liang, Xinghao Chen, Yifan Ge, Ziyi Wu, Hao Wu, Changyu Zeng, Wei Xing, Xiaoyu Shen
arXiv Computation and Language
4d ago

ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning

ToolSearcher is a reinforcement learning framework designed to improve large‑scale tool selection for large language models. It introduces category‑constrained discrimination, event‑level search modeling, and trajectory‑aligned credit allocation to better distinguish similar tools, optimize multi‑turn search, and provide fine‑grained rewards. Experiments on large‑scale benchmarks show that ToolSearcher outperforms strong baselines in iterative search and complex tool composition scenarios.

By Zhenlong Dai, Xujie Song, Zitong Wang, Tong Niu, Jian liu, Weiqiang Wang, Xiu Tang, Sai Wu, Chang Yao, Jingyuan Chen
arXiv AI
Sep 25

SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL

SLCA-GRPO addresses cross‑segment credit misattribution in tool‑calling reinforcement learning by introducing Segment‑Locked Credit Assignment (SLCA), which separates advantage estimation for tool‑invocation tokens and natural‑language summary tokens. The method leverages a Schema‑Guided LLM Simulator (SGLS) for scalable training and Hierarchical Rewards (HierR) to route execution and preference advantages appropriately. Experiments on a 7B backbone show that SLCA‑GRPO outperforms baseline methods, improving in‑domain accuracy by 2.53 pp, the Berkeley Function‑Calling Leaderboard by 1.36 pp, and $ au^2$‑Bench by 9.15 pp while reducing tool redundancy and costs.

By Yan Zhan, Shaobo Liu, Qiunan Liu, Yuanjun Shi, Siqi Xu, WeiYi Hou, Xiang Xu, Zekang Li, Weizhou Pan, Jiahong Yan
arXiv AI
Aug 26

ToolRobustBench: Stage-Wise Perturbation Evaluation and Failure Diagnosis for Tool-Calling Agents

ToolRobustBench is a stage-wise diagnostic benchmark designed to evaluate and diagnose failures in tool‑calling agents, which are large language models that select tools, provide structured arguments, and interpret tool feedback. The benchmark aligns four perturbation families—tool‑interface, user‑intent, tool‑output/observation, and runtime‑environment—with the tool‑use pipeline, attributing failures to specific stages such as tool selection, schema grounding, argument binding, and feedback handling. Experiments across 15,456 instances, 7 models, and 16 local tools reveal that while overall performance is high, robustness degrades significantly, especially under tool‑output/observation perturbations, and mixed‑family perturbations produce non‑additive failure patterns.

By YiShan Zheng, Yuan Wu, Yi Chang
Hugging Face Trending Papers
Aug 4

ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning

Historical tool-use trajectories provide valuable experience for large language model (LLM) agents to plan and coordinate tool usage. Existing approaches directly construct tool-level graphs from these trajectories, but the resulting graphs remain tied to specific tools and are hard to generalize across tool sets.

arXiv AI
Jun 16

ToolMenuBench: Benchmarking Tool-Menu Filtering Strategies for Reliable and Efficient LLM Agents

arXiv:2606. 15508v1 Announce Type: new Abstract: Tool-augmented large language model agents increasingly operate over large tool libraries, but existing evaluations often focus on whether a model can call a tool correctly rather than how the visible tool menu shapes reliability, efficiency, and safety-relevant risk exposure.

By Rahul Suresh Babu, Laxmipriya Ganesh Iyer