arXiv:2605. 03534v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) grounds answers in retrieved passages, yet relevance does not guarantee sufficiency: a topical passage may still fail to justify the answer.
By Jingxi Qiu, Zeyu Han, Cheng Huang
arXiv:2606. 23915v1 Announce Type: cross Abstract: Practice often treats automatic metrics for attribution in LLM retrieval-augmented generation as interchangeable.
By Tianyu Ding, Aditya Nannapaneni, Juan Pablo De la Cruz Weinstein
arXiv:2607. 26929v1 Announce Type: cross Abstract: The same diagnostic result can support or challenge one causal claim yet fail to address another when the claims concern different populations, outcomes, estimands, pathways, or identifying assumptions.
By Weiyi Kong, Zhuoran Li
EviScope is a new paired counterfactual benchmark that evaluates grounded language models by fixing the question while manipulating evidence—adding, removing, distracting, or contradicting it. The v1.1 dataset includes 40 four‑condition quartets with repaired counterfactual claims and span‑level support labels for automated assessment. Experiments on Qwen2.5‑7B, Llama 3.1 8B, and Gemini 3.5 Flash show that paired metrics reveal grounding behaviors hidden by simple answer accuracy, such as unsupported answers, conflict blindness, and incorrect non‑answer actions.
By Suryadeep Singh Deswal
The paper introduces TrustSwap, a counterfactual test that swaps or removes source reliability labels while keeping evidence text constant, to evaluate how retrieval‑augmented fact‑checking models respond across verdict, confidence, and search decisions. Experiments on untrained and RL‑trained models show that confidence and search largely follow labels, yet label changes can flip a significant portion of verdicts, especially in larger models. The authors propose trust‑swap augmentation (TSA) to mitigate this shortcut, demonstrating reduced verdict flip rates and maintained accuracy in several settings, though its effectiveness diminishes at larger model scales.
By Jianchang Su, Yiwei Yang, Wei Zhang
arXiv:2606. 15127v1 Announce Type: new Abstract: Reasoning models are increasingly used in settings where the final answer is not the only object of review: educational tools may show students intermediate steps, decision-support systems may require human oversight, and audit workflows may inspect traces for misleading or biased input.
By Xian Sun, Wei Gao, Yingshuo Wang, Lingdong Kong, Yanhang Li, Zhichao Fan, Zexin Zhuang, Wenlong Dong, Zhiyuan Zheng, Hrishikesh Paranjape, Abhishek Mandal, Johnny R. Zhang