The paper surveys recent advances in Retrieval Augmented Generation (RAG), a technique that integrates external retrieval into language model generation to reduce hallucinations and keep knowledge current. It introduces a four‑axis taxonomy—efficiency, defense, interactivity, and reasoning—to organize contemporary RAG research, covering retrieval methods, fusion strategies, embedding optimizations, and reinforcement learning policies. The survey also reviews evaluation practices, domain‑specific applications, and architectural variants, while highlighting ongoing challenges such as retrieval quality, reliability, domain adaptation, scalability, and explainability.
By Meghana Sunil, Shravya V, Shravan Venkatraman, Joe Dhanith PR
The paper investigates whether incorporating an evidence-support signal into retrieval evaluation for retrieval‑augmented generation (RAG) improves downstream decision‑making. Across multiple benchmarks and a TREC RAG 2025 setting, the evidence signal alters retriever rankings but its benefits vary: it does not consistently enhance retriever training, its usefulness for system selection depends on generator instructions, and it does not reliably predict answer quality on unseen topics. Human filtering of evidence‑rich passages preserves useful content, yet evaluators disagree on whether this improves final answers, indicating that evidence‑aware evaluation alone does not guarantee better downstream outcomes.
By Utshab Kumar Ghosh, Debayan Mukhopadhyay, Shubham Chatterjee
The paper introduces LOCOMO-CONV, a conversational memory benchmark that expands on the existing LoCoMo dataset with four query styles—dialog, implicit, counterfactual, and composed—designed to evaluate memory systems in realistic conversational settings. Experiments across five memory systems reveal that conversational framing uncovers significant retrieval gaps missed by traditional QA benchmarks, particularly for implicit and composed queries, and that strong retrieval does not necessarily translate into higher response quality. The study also highlights silent grounding in implicit queries, where memory enhances contextual grounding without explicitly presenting the gold fact, suggesting a need for reasoning-based memory elaboration.
By Wen-Yu Chang, Yun-Nung Chen
arXiv:2509.20377v2 Announce Type: replace-cross
Abstract: Retrieval-Augmented Generation (RAG) has significantly improved the performance of large language models (LLMs) on knowledge-intensive tasks...
By Tomoaki Isoda
arXiv:2609.23056v1 Announce Type: new
Abstract: Agentic retrieval-augmented generation (RAG) enables language models to adapt retrieval based on previously retrieved evidence, but it remains unclear...
By Kai-Hsin Chen, Wei-Yu Chen, Xuanjun Chen, Jyh-Shing Roger Jang
The paper evaluates how robust large language models (LLMs) are when using retrieval‑augmented generation (RAG) in practical settings. It investigates whether RAG always outperforms non‑RAG approaches, whether adding more retrieved documents helps, and whether the order of documents matters, using a benchmark of 1,891 samples across five datasets and three task categories. Experiments with 11 LLMs show generally high retrieval robustness, but performance varies by task and prompting strategy, indicating that adopting RAG should be considered on a case‑by‑case basis.
By Shuyang Cao, Karthik Radhakrishnan, David Rosenberg, Steven Lu, Pengxiang Cheng, Lu Wang, Shiyue Zhang
arXiv:2503. 10677v3 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) has gained significant attention in recent years for its potential to enhance natural language understanding and generation by combining large-scale retrieval systems with generative models.
By Mingyue Cheng, Yucong Luo, Jie Ouyang, Qi Liu, Huijie Liu, Li Li, Shuo Yu, Bohou Zhang, Jiawei Cao, Jie Ma, Daoyu Wang, Enhong Chen
arXiv:2606. 29090v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) has become the standard way to ground large language models in external knowledge, yet most systems retrieve a fixed number of passages for every question regardless of its difficulty.
By Ansh Kamthan
The paper "Where's Waldo? Query-language Preference under Cross-lingual Knowledge Disparities" introduces the Waldo benchmark, a multilingual QA dataset built from Wikipedia that focuses on knowledge gaps and conflicts across languages. It evaluates eight models in five languages and finds that when a fact is missing in one language, models tend to use evidence from the other language, but when conflicting accounts exist, responses align strongly with the query language, leading to different answers for semantically identical questions. The study also explores mitigation strategies, including ablating attention heads and LoRA-based training, which can reduce the preference gap by up to 61.5%.
By Dayeon Ki, Ruochen Zhang, Silviu Cucerzan, Ryen W. White, Ning Gao
arXiv:2607. 11267v1 Announce Type: cross Abstract: In the rapidly evolving landscape of information retrieval systems, the ability to adapt and improve through user feedback is paramount.
By Tatiana Pelc, Gila Kamhi, Asaf Avrahamy, Adi Fledel-Alon
Retrieval-Augmented Generation (RAG) has become the standard way to ground large language models in external knowledge, yet most systems retrieve a fixed number of passages for every question regardless of its difficulty. This wastes computation on easy questions, starves hard ones, and gives no signal for when a generated answer can be trusted.
arXiv:2609.36700v1 Announce Type: cross
Abstract: When conversing with large language models (LLMs), users often begin with a simple question and build towards a multi-hop question through follow-up...
By Pranav Handa, Ariful Azad