arXiv AI By Jo\~ao Coelho, Jo\~ao Magalh\~aes, Bruno Martins, Chenyan Xiong

Effective Reinforcement Learning for Agentic Search by Recycling Zero-Variance Queries During Training

Read the original on arXiv AI →

arXiv:2606. 10709v1 Announce Type: cross Abstract: The use of GRPO-style algorithms has become the standard strategy for training LLM search agents under outcome-only rewards.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 25

Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization

arXiv:2608.23311v1 Announce Type: new Abstract: Policy optimization (PO) for Large Language Models faces a stability--exploration trade-off, currently mediated by an action-side Policy-KL regularizer...

By Xianlei Zhou, Xiangdi Meng, Yu He, Tianyu Qi, Shuyan Guan, Xianli Zhang, Jian Zhang, Xin Li, Qika Lin, Jun Liu
arXiv AI
Sep 2

One Policy, Any Budget: Internalizing Budget-Aware Search via Reinforcement Learning

The paper introduces AnySearch, a reinforcement‑learning framework that trains a single policy to perform budget‑aware search for large language models under any budget constraint. The training proceeds in two phases: first, the agent learns with explicit budget state injection and structured reasoning prompts under linearly decaying budgets; second, the scaffold is removed and the agent adapts to randomly sampled budgets that match deployment conditions. The reward combines answer accuracy and budget efficiency, with adaptive weighting to emphasize efficiency for high‑accuracy queries and reduce it for low‑accuracy ones. Experiments on seven QA benchmarks demonstrate that AnySearch outperforms baselines across all budget scales, generalizes to unseen constraints, and improves tool productivity without excessive token overhead.

By Xiaowei Sun, Jin Li, Yili Hong, Yikun Fu, Yanghua Xiao