arXiv:2605. 03534v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) grounds answers in retrieved passages, yet relevance does not guarantee sufficiency: a topical passage may still fail to justify the answer.
By Jingxi Qiu, Zeyu Han, Cheng Huang
arXiv:2606. 23695v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) grounds Large Language Models in external knowledge, yet current evaluations rely on discrete heuristics that suffer from ''epistemic blindness'' - failing to distinguish genuine contextual information extraction from parametric memory recall.
By Barak Or
arXiv:2608.29307v1 Announce Type: cross
Abstract: Language models increasingly answer questions by consulting retrieved documents rather than memory alone, a design now common in search assistants an...
By Sai Krishna Reddy Mulakkayala, Niki van Stein, Aske Plaat
The paper introduces TRACE, a fine‑tuning framework for Retrieval‑Augmented Generation (RAG) that addresses conflicts between retrieved knowledge and a model’s internal knowledge. TRACE uses multi‑agent debate traces to identify correct and incorrect candidates and answer‑shift patterns, providing fine‑grained supervision for reliable knowledge‑source selection. It also incorporates an answer‑completeness regularization mechanism to prevent empty, overly short, or prematurely terminated responses, thereby improving robustness against misleading retrieved content and enhancing answer quality.
By Zhengchen Huang, Yundong Sun, Minrui Song, Shuanglong Yao, Ye Liu, Ji Chen, Xing Wang
The paper investigates whether incorporating an evidence-support signal into retrieval evaluation for retrieval‑augmented generation (RAG) improves downstream decision‑making. Across multiple benchmarks and a TREC RAG 2025 setting, the evidence signal alters retriever rankings but its benefits vary: it does not consistently enhance retriever training, its usefulness for system selection depends on generator instructions, and it does not reliably predict answer quality on unseen topics. Human filtering of evidence‑rich passages preserves useful content, yet evaluators disagree on whether this improves final answers, indicating that evidence‑aware evaluation alone does not guarantee better downstream outcomes.
By Utshab Kumar Ghosh, Debayan Mukhopadhyay, Shubham Chatterjee
CARGO is a framework for evaluating agentic AI systems in production that addresses the problem of reference-instance divergence (RID), where reference-based judges penalize correct answers that involve different entity identifiers. It treats retrieved references as procedural exemplars, grounds judgments in the live instance’s context, assigns a three-way status to claims, and gates evaluation by retrieval confidence. Using the CARGO-Bench diagnostic suite, CARGO eliminates false penalties and improves discrimination while revealing a limitation in detecting procedural corruptions.
By Mukul Chhabra, Shail Patel, Luigi Medrano
arXiv:2512. 11614v3 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) relies on retrieved context to guide large language models (LLM), yet treats the retrieval as a heuristic rather than verifiable evidence -- leading to unsupported answers, hallucinations, and reliance on spurious context.
By Bj\"orn Deiseroth, Max Henning H\"oth, Kristian Kersting, Letitia Parcalabescu
The paper investigates how memory systems can answer a current query correctly yet fail to retain distinctions needed for later updates. Using a paired‑history audit, the authors evaluate 24 history pairs across six synthetic mechanisms and two model backends, achieving perfect reveal accuracy on DeepSeek and high accuracy on GLM. Record‑level audits reveal specific failures in structured reveal memories and frontier late‑reference adequacy, and the authors test a label‑equivariant repair that only partially restores correctness.
By Guangzhe Zhang
The paper introduces "stale‑document poisoning," a temporal alignment failure where outdated retrieval evidence causes retrieval‑augmented generation models to produce incorrect answers even when the model could answer correctly without retrieval. A benchmark of 317 verified knowledge reversals across medicine, law, software, and platform policy shows that outdated evidence flips 30–37% of Llama and Qwen answers, rising to 66–75% when models are explicitly instructed to trust the document. The study demonstrates that providing explicit validity information and a recency‑aware re‑ranker can substantially reduce poisoning, highlighting the need for models to assess whether retrieved evidence remains applicable.
By Md Shamim Ahmed, Lukas Galke Poech, Richard R\"{o}ttger
arXiv:2606. 01120v1 Announce Type: new Abstract: In RAG-based fact-checking, LLMs are increasingly used as verifiers to check given claims against retrieved evidence.
By Yuxi Sun, Wenbo Shang, Wei Gao, Xin Huang, Jing Ma
arXiv:2601. 19827v4 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) extends large language models (LLMs) beyond parametric knowledge, yet it is unclear when iterative retrieval-reasoning loops meaningfully outperform static RAG, particularly in scientific domains with multi-hop reasoning, sparse domain knowledge, and heterogeneous evidence.
By Mahdi Astaraki, Mohammad Arshi Saloot, Ali Shiraee Kasmaee, Hamidreza Mahyar, Soheila Samiee
arXiv:2609.37469v1 Announce Type: cross
Abstract: Retrieval-augmented generation (RAG) grounds large language models in external sources, but retrieved passages often name the right entities without...
By Suting Chen, Peichun Hua, Yunming Xiao