arXiv:2605. 03534v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) grounds answers in retrieved passages, yet relevance does not guarantee sufficiency: a topical passage may still fail to justify the answer.
By Jingxi Qiu, Zeyu Han, Cheng Huang
The paper investigates how memory systems can answer a current query correctly yet fail to retain distinctions needed for later updates. Using a paired‑history audit, the authors evaluate 24 history pairs across six synthetic mechanisms and two model backends, achieving perfect reveal accuracy on DeepSeek and high accuracy on GLM. Record‑level audits reveal specific failures in structured reveal memories and frontier late‑reference adequacy, and the authors test a label‑equivariant repair that only partially restores correctness.
By Guangzhe Zhang
arXiv:2608.31016v1 Announce Type: cross
Abstract: Ambient AI scribes draft clinical notes, and published audits find their dominant error is omission: information the encounter established that the n...
By Sebastian Fox, Luke Markham, Ryan Lail, Michael Karotsieris
arXiv:2607. 28871v1 Announce Type: cross Abstract: When a repair agent runs a test and sees it pass, the result is treated as evidence about the reported defect.
By Xiaonan Xu, Wenjing Wu
arXiv:2609.38021v1 Announce Type: cross
Abstract: We evaluate an auditable long-term memory system on LongMemEval-S. Its retrieval chain uses hybrid candidate retrieval, cross-encoder reranking, cove...
By Christopher J. Chanhnourack
The paper introduces AgenticRAG-FP, an interventional benchmark designed to attribute causal failures in agentic retrieval‑augmented generation (RAG) systems. By injecting a certified fault at a specified hop and re‑executing the downstream trajectory, the benchmark evaluates whether post‑hoc diagnostics can correctly identify the fault’s location. Experiments on MuSiQue questions show that coverage‑based diagnosis performs well at hop 1 but poorly at later hops, while counterfactual probes reveal varying diagnostic success depending on propagation depth.
By Lauren Pothuru
arXiv:2609.08279v1 Announce Type: cross
Abstract: Agent memory systems must discard stored information when their history exceeds a fixed token budget. Existing budget-accuracy frontiers quantify the...
By Chen Shen
The study investigates how language models equipped with tools can still produce unsupported final claims, even when a single tool call could resolve the uncertainty. It defines two metrics—occurrence (how often unsupported claims arise) and conditional repair (how often they are fixed when evidence is provided). Experiments on Qwen3-32B and Gemma 4 show that providing the missing evidence consistently repairs all unsupported claims in the Qwen3-32B setup, while the Gemma 4 model never produced unsupported claims under the tested conditions.
By Justin Bronder
arXiv:2606. 17062v1 Announce Type: cross Abstract: Radiology report evaluation must distinguish clinical compatibility from surface similarity, because negation, laterality, or normal-abnormal polarity can reverse a finding.
By Zhenhong Yang, Zhuoyun Liu, Jintao Fei, Wen Tang, Shichao Quan, Jun Zhao, Jun Xu
arXiv:2607. 26929v1 Announce Type: cross Abstract: The same diagnostic result can support or challenge one causal claim yet fail to address another when the claims concern different populations, outcomes, estimands, pathways, or identifying assumptions.
By Weiyi Kong, Zhuoran Li
The paper introduces ClaimReceipt, a specification and verifier that checks whether a claim in an agent evaluation can be recomputed from retained evidence (sufficiency) and whether the evidence covers the entire experiment set (coverage). Using the CR‑2 verifier on 1,392 historical records, the authors demonstrate accurate reproduction of audit verdicts, non‑redundant field groups, and zero false positives on semantic faults. In a prospective CR‑3 run, the system correctly flags missing receipts and preserves coverage when private evidence is withheld, while adding minimal overhead to inference time and transaction size.
By Peiying Zhu, Sidi Chang
The paper introduces "stale‑document poisoning," a temporal alignment failure where outdated retrieval evidence causes retrieval‑augmented generation models to produce incorrect answers even when the model could answer correctly without retrieval. A benchmark of 317 verified knowledge reversals across medicine, law, software, and platform policy shows that outdated evidence flips 30–37% of Llama and Qwen answers, rising to 66–75% when models are explicitly instructed to trust the document. The study demonstrates that providing explicit validity information and a recency‑aware re‑ranker can substantially reduce poisoning, highlighting the need for models to assess whether retrieved evidence remains applicable.
By Md Shamim Ahmed, Lukas Galke Poech, Richard R\"{o}ttger