arXiv AI By Jo\~ao Coelho, Jo\~ao Magalh\~aes, Bruno Martins, Chenyan Xiong

Effective Reinforcement Learning for Agentic Search by Recycling Zero-Variance Queries During Training

Read the original on arXiv AI →

arXiv:2606. 10709v1 Announce Type: cross Abstract: The use of GRPO-style algorithms has become the standard strategy for training LLM search agents under outcome-only rewards.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.