arXiv:2609.23056v1 Announce Type: new
Abstract: Agentic retrieval-augmented generation (RAG) enables language models to adapt retrieval based on previously retrieved evidence, but it remains unclear...
By Kai-Hsin Chen, Wei-Yu Chen, Xuanjun Chen, Jyh-Shing Roger Jang
The paper introduces RAVEL, a retrieval‑aware online reinforcement learning framework designed to improve interactive retrieval under partial evidence. RAVEL begins with supervised question generation, directly observes the top‑4 retrieval candidates, and refines its question policy using rank feedback from the full question‑answer‑retrieval loop. Experiments on the Interactive‑PEDES dataset demonstrate that RAVEL progressively enhances retrieval performance over five interaction rounds, reallocating questioning toward localized open‑ended attributes that yield the greatest gains on challenging queries.
By Lyucheng Qian, John Yuehan Zhang, Pingyu Wang
arXiv:2606.13120v2 Announce Type: replace
Abstract: Search Agents -- large language models augmented with search tools -- have intensified the need for future-proof evaluation benchmarks. Existing be...
By Yunhan Wang, Jiaan Wang, Lianzhe Huang, Xianfeng Zeng, Fandong Meng
arXiv:2608.22266v1 Announce Type: new
Abstract: In the context of information seeking, conversational agents are undergoing an evolution from reactive tools to proactive, personalized assistants. A c...
By Zhihong Cao, Chen Huang
HintMiner is an automatic tool that mines question hints from web Q&A posts using a language‑model‑based MiningNet. It retrieves many Q&A posts, extracts hints via a transformer‑based encoder‑decoder with copying mechanisms, and is trained with a self‑supervised objective on large online data. Evaluated on 60,000 Stack Overflow questions, HintMiner achieves an average BLEU score of 36.17% and ROUGE‑2 of 36.29%.
By Zhenyu Zhang, JiuDong Yang
arXiv:2504. 07385v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) become increasingly used for question-answering (QA), relying on static, pre-annotated references for evaluation poses significant challenges in cost, scalability, and completeness.
By Sher Badshah, Ali Emami, Hassan Sajjad
arXiv:2609.07093v2 Announce Type: replace
Abstract: Retrieval-augmented generation (RAG) enables large language models (LLMs) to answer questions by accessing external knowledge and has been widely a...
By Yifan Wang, Xinkui Lin, Yongxiu Xu, Shen Gao, Ruochen Yang, Kun Huang, Yubin Wang, Jie Wu, Wei Liu, Jian Luan, Hongbo Xu, Shuo Shang
The paper presents a tri‑agent framework for evaluating large language models’ question‑clarification abilities. It involves a Question Clarifying Agent that identifies ambiguities and asks follow‑up questions, a Respondent Agent that simulates human replies, and an Evaluator Agent that judges the dialogue using metrics such as ambiguity handling, question quality, dialogue efficiency, language appropriateness, and intent alignment. The authors illustrate the approach with synthetic supply‑chain data and discuss validating the evaluator against human judgments.
By Yikai Zhao, Saurabh Pandey, Pradeep Kumar Misra
Large Language Models (LLMs) are increasingly deployed in interactive systems where understanding user intent precisely is paramount. A key capability for such systems is effective question clarificat...
arXiv:2608. 09934v1 Announce Type: cross Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors.
By Vitalii Belov, Artyom Sosedka, Andrey Sakhovskiy, Elizaveta Kovtun, Artyom Boyarskikh, Semen Budennyy
The paper presents a system for AI contact centers that answers questions only from a closed set of verified QA units, returning the unit verbatim or routing to clarification, abstention, or handoff. The index is enriched offline using staged linguistic seeding (SLS), where human-authored slot recipes are expanded by GPT‑4.1‑mini and lightly filtered by humans, enabling a single retrieval pass without query-time generation. On held‑out data from two industrial domains, SLS improves hybrid retrieval recall at rank 1 to 0.881/0.930 and outperforms doc2query by 0.20/0.32, while also reducing unsupported content from 7‑13% to near 0%.
By Hyeonseop Yoon, Jeong-Eun Park
arXiv:2608.21558v1 Announce Type: cross
Abstract: Recent advances in LLMs and the adoption of RAG systems in industry have created a need for domain-specific question-answer datasets that can assess...
By Lorenz Brehme, Adam Jatowt