arXiv AI

CoBRA: Learning Tool-Use Boundaries via Counterfactual Margins

CoBRA is a counterfactual boundary‑learning framework designed to improve when a tool‑augmented language model should call an external tool. It builds internal and external experts from the same base model, collects paired trajectories, and estimates the reward margin between answering with and without tools. Using these margins, CoBRA partitions data into internal‑favored, external‑favored, and ambiguous cases, then applies Boundary‑Aware Cold‑Start SFT and MARS‑RL to optimize boundary decisions, leading to more efficient tool use and better accuracy on tool‑dependent out‑of‑distribution questions.

arXiv Computation and Language
3d ago

ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning

ToolSearcher is a reinforcement learning framework designed to improve large‑scale tool selection for large language models. It introduces category‑constrained discrimination, event‑level search modeling, and trajectory‑aligned credit allocation to better distinguish similar tools, optimize multi‑turn search, and provide fine‑grained rewards. Experiments on large‑scale benchmarks show that ToolSearcher outperforms strong baselines in iterative search and complex tool composition scenarios.

By Zhenlong Dai, Xujie Song, Zitong Wang, Tong Niu, Jian liu, Weiqiang Wang, Xiu Tang, Sai Wu, Chang Yao, Jingyuan Chen
arXiv Computation and Language
Sep 16

Spurious Tool Use: When RL Agents Learn the Wrong Reason to Act

The paper investigates how reinforcement learning can cause large language model agents to adopt shortcut policies for tool use, relying on superficial prompt cues rather than actual task needs. By creating synthetic environments that mix factual QA and math reasoning, the authors show that agents often invoke tools when cues are present, even when those tools are unnecessary, with spurious invocation rates rising up to 39%. They find that shortcut learning occurs mainly when agents have already mastered the target tool and that semantic alignment between cues and tools amplifies the effect. To counter this, they propose a dense, decision-level reward where an LLM judge assesses tool necessity, which reduces cue-driven tool use while maintaining performance.

By Yiwei Yang, Haoxiang Zhang, Bingbing Wen, Yao Lu, Yuchen Wu, Lei Zhang, Julian McAuley, Pan Lu, Bill Howe
arXiv Computation and Language
Aug 25

GTA-RAG: Graph-Trajectory-Augmented Reinforcement Learning for Multi-Turn Retrieval-Augmented Reasoning

arXiv:2608.22479v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) enables LLMs to access external knowledge for answering knowledge-intensive questions. For complex multi-hop quest...

By Jun Chen, Yongchao Liu, Pengyu Qiu, Jiajun Zheng, Juelu Zhang, Yujie Zeng, Qin Zhang, Ziyue Qiao, Xiao Luo
arXiv AI
Sep 18

MATCH: Model-Aware Tool Learning with Curriculum Scheduling and Hierarchically Gated Rewards

MATCH is a closed‑loop framework for model‑aware tool learning that combines curriculum scheduling with hierarchically gated rewards. It introduces Model‑Aware Curriculum Learning (MACL), which dynamically adjusts sample difficulty based on reward signals, and Hierarchical Tool‑call Gated Reward (HTGR), which allocates credit at the tool name, argument key, and argument value levels only when prerequisites are met. Experiments on API‑Bank and BFCL V3 show MATCH achieving 72.19% and 62.87% overall accuracy, outperforming both supervised and RL‑based baselines across multiple backbone models.

By Shihao Liu, Hao Yin, Lijun Liu, Zhengzong Chen, Yuanyuan Zhao, Fei Huang
arXiv AI
Aug 26

Joint Optimization of Tool Creation and Use for Large Language Model Agents

The paper introduces SMITH, a reinforcement learning framework that jointly trains large language models to create and use tools within a single policy. By alternating between build and use tasks and employing separate reward signals for schema, code, and outcome failures, SMITH enables a 4B Qwen3 model to achieve state‑of‑the‑art accuracy on procedural reasoning benchmarks, outperforming larger untrained models and improving performance on downstream tasks when its tools are applied.

By Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen, Shao-Hua Sun, Hung-yi Lee
arXiv AI
Jul 21

From Evidence to Trajectory: Abductive Reasoning Path Synthesis for Retrieval-Augmented Generation Agents Development

arXiv:2509. 23071v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) agent development is hindered by the lack of executable ground-truth agent-environment interaction trajectories.

By Muzhi Li, Jinhu Qi, Yihong Wu, Minghao Zhao, Liheng Ma, Yifan Li, Xinyu Wang, Zhenghan Tai, Zixing Song, Yingxue Zhang, Ho-fung Leung, Irwin King