arXiv:2608. 08944v1 Announce Type: cross Abstract: A failed retrieval-augmented generation (RAG) answer can be consistent with several unseen responses to evidence repair.
By Wenzhang Du
The paper introduces a penalty‑aware evaluation framework for Retrieval‑Augmented Generation (RAG) systems that uses asymmetric scoring, knowledge‑gap canaries, and a failure‑attribution pipeline. Applying this framework to three commercial RAG products and a baseline on SimpleQA‑Verified, the authors find that while overall accuracy is similar across systems, canary violation rates vary dramatically, showing that systems differ more in when they answer than in what they answer. The study demonstrates that penalty‑aware scoring can reorder system rankings and is robust across different penalty settings.
By Alden Do Rosario, Hussein Younes, Felipe Pires
DynaKRAG is a unified framework that learns a state‑conditioned policy to control evidence acquisition in multi‑hop retrieval‑augmented generation. It uses a deterministic validity layer to build an action set, a learned continuation gate to decide between generating an answer or gathering more evidence, and an advantage scorer to rank evidence operations by predicted gain. Across HotpotQA, 2Wiki, and MuSiQue with various backbone models, DynaKRAG achieves top EM and F1 scores, improves token and retrieval efficiency, and enables terminal evidence compression that reduces context size while boosting answer quality.
By Chenyu Zhou, Yaqi Wu, Xiaolei Guo, Jiaqi Huang, Xianfa Zhang, Junxu Zhang, Zhuo Yu, Zhubo Shi, Jianghao Lin, Dongdong Ge
arXiv:2605. 14473v4 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) is usually evaluated by whether the final answer is correct.
By Yihang Chen, Pin Qian, Su Wang, Sipeng Zhang, Huan Xu, Shuhuai Lin, Xinpeng Wei
arXiv:2609.38021v1 Announce Type: cross
Abstract: We evaluate an auditable long-term memory system on LongMemEval-S. Its retrieval chain uses hybrid candidate retrieval, cross-encoder reranking, cove...
By Christopher J. Chanhnourack
arXiv:2608.31016v1 Announce Type: cross
Abstract: Ambient AI scribes draft clinical notes, and published audits find their dominant error is omission: information the encounter established that the n...
By Sebastian Fox, Luke Markham, Ryan Lail, Michael Karotsieris