arXiv Computation and Language By Tzu-Han Lin, Wei-Lin Chen, Chen-An Li, Hung-yi Lee, Yun-Nung Chen, Yu Meng

AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning

Read the original on arXiv Computation and Language →

AdaSearch introduces a two‑stage reinforcement learning framework that separates problem solving from the decision to search in large language models. By using an F1‑based decision metric, it explicitly evaluates when external search is needed, reducing unnecessary search calls while maintaining high question‑answering performance. Experiments show that AdaSearch improves search‑decision quality with only a minor impact on accuracy compared to always‑search strategies.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

Hugging Face Trending Papers
Jul 12

To Answer or to Abstain: Mitigating Search-Agent Hallucinations via Abstention-Aware Reinforcement Learning

Recent advances in equipping Large Language Models (LLMs) with search tools and outcome-reward reinforcement learning (RL) have achieved new state-of-the-art results on open-domain QA tasks. However, we argue that current training paradigms harbor a critical vulnerability: they predominantly reward correct answers but fail to penalize fabricated ones when retrieval fails, thereby implicitly exacerbating hallucinations.

arXiv AI
Jul 14

To Answer or to Abstain: Mitigating Search-Agent Hallucinations via Abstention-Aware Reinforcement Learning

arXiv:2607. 10738v1 Announce Type: cross Abstract: Recent advances in equipping Large Language Models (LLMs) with search tools and outcome-reward reinforcement learning (RL) have achieved new state-of-the-art results on open-domain QA tasks.

By Fengji Zhang, Tianyu Fan, Yuxiang Zheng, Xinyao Niu, Chengen Huang, Jacky Keung, Bei Chen
arXiv AI
1d ago

BRIDGE: Bilevel Retrieval-Credit-Aware Agentic Reinforcement Learning

The paper introduces BRIDGE, a bilevel optimization framework that jointly trains a large language model (LLM) and a retriever for agentic reinforcement learning (ARL). It demonstrates that adapting the retriever before the policy yields better rewards, and that BRIDGE outperforms existing methods on seven open‑domain QA benchmarks and medical QA tasks, achieving significant gains in accuracy and reasoning quality.

By Quan Xiao, Mingda Liu, Gaowen Liu, Katsuki Fujisawa, Tianyi Chen
arXiv AI
6d ago

Search-Aware Reinforcement Learning for Multi-Component Query Understanding in Roblox Game Search

The paper introduces a search‑aware reinforcement learning framework for multi‑component query understanding in Roblox game search. It first uses teacher‑student supervised fine‑tuning to create a schema‑compliant policy, then applies reinforcement learning that optimizes each query‑understanding component with component‑specific rewards derived from live search engine interactions. Experiments show that this approach improves per‑component utility and overall search quality, raising NDCG@20 by 8.9 points over the supervised baseline and 3.5 points over a single end‑to‑end reward strategy.

By Nayoung Choi, Shengjian Chen, Xiaokai Wei, Wenzheng Zhang, Daiyao Yi, Rachit Pareek, Vincent Su, Michelle Gong, Jinho D. Choi