arXiv AI By Wenhao Zou, Xianglong Liu, Wendong Bi, Hanjie Wang, Simin Zhao, Gong Zhi

CoBRA: Learning Tool-Use Boundaries via Counterfactual Margins

Read the original on arXiv AI →

CoBRA is a counterfactual boundary‑learning framework designed to improve when a tool‑augmented language model should call an external tool. It builds internal and external experts from the same base model, collects paired trajectories, and estimates the reward margin between answering with and without tools. Using these margins, CoBRA partitions data into internal‑favored, external‑favored, and ambiguous cases, then applies Boundary‑Aware Cold‑Start SFT and MARS‑RL to optimize boundary decisions, leading to more efficient tool use and better accuracy on tool‑dependent out‑of‑distribution questions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
3d ago

ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning

ToolSearcher is a reinforcement learning framework designed to improve large‑scale tool selection for large language models. It introduces category‑constrained discrimination, event‑level search modeling, and trajectory‑aligned credit allocation to better distinguish similar tools, optimize multi‑turn search, and provide fine‑grained rewards. Experiments on large‑scale benchmarks show that ToolSearcher outperforms strong baselines in iterative search and complex tool composition scenarios.

By Zhenlong Dai, Xujie Song, Zitong Wang, Tong Niu, Jian liu, Weiqiang Wang, Xiu Tang, Sai Wu, Chang Yao, Jingyuan Chen
arXiv Computation and Language
Sep 16

Spurious Tool Use: When RL Agents Learn the Wrong Reason to Act

The paper investigates how reinforcement learning can cause large language model agents to adopt shortcut policies for tool use, relying on superficial prompt cues rather than actual task needs. By creating synthetic environments that mix factual QA and math reasoning, the authors show that agents often invoke tools when cues are present, even when those tools are unnecessary, with spurious invocation rates rising up to 39%. They find that shortcut learning occurs mainly when agents have already mastered the target tool and that semantic alignment between cues and tools amplifies the effect. To counter this, they propose a dense, decision-level reward where an LLM judge assesses tool necessity, which reduces cue-driven tool use while maintaining performance.

By Yiwei Yang, Haoxiang Zhang, Bingbing Wen, Yao Lu, Yuchen Wu, Lei Zhang, Julian McAuley, Pan Lu, Bill Howe
arXiv Computation and Language
Aug 25

GTA-RAG: Graph-Trajectory-Augmented Reinforcement Learning for Multi-Turn Retrieval-Augmented Reasoning

arXiv:2608.22479v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) enables LLMs to access external knowledge for answering knowledge-intensive questions. For complex multi-hop quest...

By Jun Chen, Yongchao Liu, Pengyu Qiu, Jiajun Zheng, Juelu Zhang, Yujie Zeng, Qin Zhang, Ziyue Qiao, Xiao Luo
arXiv AI
Sep 18

MATCH: Model-Aware Tool Learning with Curriculum Scheduling and Hierarchically Gated Rewards

MATCH is a closed‑loop framework for model‑aware tool learning that combines curriculum scheduling with hierarchically gated rewards. It introduces Model‑Aware Curriculum Learning (MACL), which dynamically adjusts sample difficulty based on reward signals, and Hierarchical Tool‑call Gated Reward (HTGR), which allocates credit at the tool name, argument key, and argument value levels only when prerequisites are met. Experiments on API‑Bank and BFCL V3 show MATCH achieving 72.19% and 62.87% overall accuracy, outperforming both supervised and RL‑based baselines across multiple backbone models.

By Shihao Liu, Hao Yin, Lijun Liu, Zhengzong Chen, Yuanyuan Zhao, Fei Huang