arXiv AI

SAAS: Self-Aware Reinforcement Learning for Over-Search Mitigation in Agentic Search

arXiv:2605. 29796v3 Announce Type: replace Abstract: Agentic search enables LLMs to solve complex multi-hop questions through iterative reasoning and external search.

arXiv Computation and Language
4d ago

Traverse: Learning When to Remember, Reset, and Redirect for Long-Horizon Web Search

The paper introduces Traverse, an autonomous web‑search agent that manages its search process through three states—Rubric, Answer, and Verify—while using a Seal Memory tool for active context management. Reinforcement learning is employed to train the agent, but a training instability called Seal Collapse is mitigated by training only the final segment after context management. The resulting 35B model achieves state‑of‑the‑art performance on BrowseComp and related benchmarks, outperforming comparable open‑source systems.

By Jingyuan Ma, Lynx Aster, He Zhang, Siyao Song, Weijie Yuan, Zhe Zhang, Kai Jia, Zhifang Sui
arXiv Computation and Language
Sep 2

AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning

AdaSearch introduces a two‑stage reinforcement learning framework that separates problem solving from the decision to search in large language models. By using an F1‑based decision metric, it explicitly evaluates when external search is needed, reducing unnecessary search calls while maintaining high question‑answering performance. Experiments show that AdaSearch improves search‑decision quality with only a minor impact on accuracy compared to always‑search strategies.

By Tzu-Han Lin, Wei-Lin Chen, Chen-An Li, Hung-yi Lee, Yun-Nung Chen, Yu Meng
arXiv AI
Jun 2

Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses

arXiv:2606. 02373v1 Announce Type: new Abstract: Search agents are often trained as policies over growing transcripts: the model must decide how to search while also remembering what it has seen, which evidence is useful, which constraints remain open, and which claims have actually been checked.

By Pengcheng Jiang, Zhiyi Shi, Kelly Hong, Xueqiang Xu, Jiashuo Sun, Jimeng Sun, Hammad Bashir, Jiawei Han
arXiv AI
Jul 14

To Answer or to Abstain: Mitigating Search-Agent Hallucinations via Abstention-Aware Reinforcement Learning

arXiv:2607. 10738v1 Announce Type: cross Abstract: Recent advances in equipping Large Language Models (LLMs) with search tools and outcome-reward reinforcement learning (RL) have achieved new state-of-the-art results on open-domain QA tasks.

By Fengji Zhang, Tianyu Fan, Yuxiang Zheng, Xinyao Niu, Chengen Huang, Jacky Keung, Bei Chen
arXiv AI
Jun 30

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization

arXiv:2602. 11351v2 Announce Type: replace Abstract: Proactive large language model (LLM) agents aim to actively plan, query, and interact over multiple turns, enabling efficient task completion beyond passive instruction following and making them essential for real-world, user-centric applications.

By Yihang Yao, Zhepeng Cen, Haohong Lin, Shiqi Liu, Zuxin Liu, Jiacheng Zhu, Zhang-Wei Hong, Laixi Shi, Ding Zhao