The paper introduces Evidence Sufficiency Boundary Training, a framework that teaches models to abstain from answering until the supplied evidence is fully sufficient, and to remain stable when additional redundant evidence is added. By constructing ordered evidence chains from datasets such as HotpotQA, 2WikiMultiHopQA, and MuSiQue, the method applies level supervision, a boundary flip margin, post‑boundary stability, and answer recall protection. Experiments with Qwen2.5‑3B‑Instruct and LoRA adaptation show improved boundary localization (flip accuracy 0.807 vs 0.781) and a lower unsupported‑answer rate (0.095 vs 0.101) while maintaining competitive raw QA F1.
By Haruto Sato, Yuki Tanaka, Ren Nakamura, Aoi Kobayashi, Mei Ito
Grounded question answering systems should answer only when the supplied evidence supports the answer. In multi-hop QA, this requirement is difficult because partial evidence can make an unsupported a...
arXiv:2609.37469v1 Announce Type: cross
Abstract: Retrieval-augmented generation (RAG) grounds large language models in external sources, but retrieved passages often name the right entities without...
By Suting Chen, Peichun Hua, Yunming Xiao
arXiv:2606. 06197v1 Announce Type: cross Abstract: Question answering (QA) systems have achieved notable progress with the advent of large language models (LLMs).
By Hafez Abdelghaffar, Ahmed Alansary, Ali Hamdi
arXiv:2609.24799v1 Announce Type: new
Abstract: Post-training quantization (PTQ) enables efficient deployment of large language models, and PTQ methods are usually optimized and evaluated with generi...
By Yeji Kim, Mi-Young Kim, Randy Goebel
arXiv:2501. 06286v2 Announce Type: replace-cross Abstract: Multi-hop question answering requires a system to identify and integrate evidence distributed across documents, yet large language models remain vulnerable to irrelevant context.
By Iman Barati, Arash Ghafouri, Behrouz Minaei-Bidgoli
arXiv:2606. 29090v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) has become the standard way to ground large language models in external knowledge, yet most systems retrieve a fixed number of passages for every question regardless of its difficulty.
By Ansh Kamthan
The paper investigates whether incorporating an evidence-support signal into retrieval evaluation for retrieval‑augmented generation (RAG) improves downstream decision‑making. Across multiple benchmarks and a TREC RAG 2025 setting, the evidence signal alters retriever rankings but its benefits vary: it does not consistently enhance retriever training, its usefulness for system selection depends on generator instructions, and it does not reliably predict answer quality on unseen topics. Human filtering of evidence‑rich passages preserves useful content, yet evaluators disagree on whether this improves final answers, indicating that evidence‑aware evaluation alone does not guarantee better downstream outcomes.
By Utshab Kumar Ghosh, Debayan Mukhopadhyay, Shubham Chatterjee
The paper introduces TRACE, a fine‑tuning framework for Retrieval‑Augmented Generation (RAG) that addresses conflicts between retrieved knowledge and a model’s internal knowledge. TRACE uses multi‑agent debate traces to identify correct and incorrect candidates and answer‑shift patterns, providing fine‑grained supervision for reliable knowledge‑source selection. It also incorporates an answer‑completeness regularization mechanism to prevent empty, overly short, or prematurely terminated responses, thereby improving robustness against misleading retrieved content and enhancing answer quality.
By Zhengchen Huang, Yundong Sun, Minrui Song, Shuanglong Yao, Ye Liu, Ji Chen, Xing Wang
The paper investigates where and how large language models encode signals that indicate answer correctness. By examining hidden states, token probabilities, residual-stream features, attention, and their combinations, the authors find that correctness signals are concentrated in the answer span and that different signal families complement each other. Fusing these signals improves robustness, especially under distribution shifts, and can be used to control retrieval in downstream tasks.
By Yuansen Liu, Yixuan Tang, Anthony Kum Hoe Tung
arXiv:2607. 04223v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) reduces but does not eliminate hallucination, and existing detectors return a single answer-level score that does not indicate which sentence is unsupported, or why.
By Mohamed Aly Bouke
arXiv:2607. 10491v1 Announce Type: new Abstract: Retrieval-augmented generation grounds large language models in external evidence, but most pipelines still treat retrieved passages as deterministic and mutually consistent context.
By S M Asif Hossain, Ruksat Khan Shayoni, M. F. Mridha