arXiv:2606. 26449v1 Announce Type: cross Abstract: Retrieval-augmented systems routinely present citations alongside generated answers, yet a citation does not confirm that the corresponding source meaningfully shaped the output.
By Mohammad Faizan, Dalal Alharthi
arXiv:2609.15830v1 Announce Type: cross
Abstract: Retrieval-augmented generation (RAG) can improve access to complex information; however, retrieving evidence alone does not ensure that answers are g...
By Sumit Barua, Guan Hong, Halil Dursunoglu, Charles Rodgers, Alvis Fong
The paper introduces a claim‑gated audit framework for generative search, ensuring that a query, source, and answer tuple is only considered resolved when relationship evidence, answer adoption, materiality, and disclosure are all present. It distinguishes this audit endpoint from citation support and review priority, tying decisions to versioned evidence spans and implementing a reference checker to enforce the contract. Experiments on a synthetic dataset confirm that the system correctly handles all 81 predicate combinations and rejects 192 malformed records, while ablation studies isolate endpoint logic from missing‑evidence handling.
By Kainan Zhou, Chuhong Xu, Gangzhen Qian, Zhaoyi Li
arXiv:2607. 22584v1 Announce Type: new Abstract: Standard Retrieval-Augmented Generation pipelines rank retrieved documents by semantic similarity alone, without accounting for source provenance or credibility.
By Yuktha Tata Koganti, Hugo Garrido-Lestache Belinchon
arXiv:2607. 17883v1 Announce Type: cross Abstract: Enterprises will not deploy AI agents they cannot trust, and the most-cited reason for distrust is hallucination: confident, fluent output that is simply not true.
By Bogdan Raduta, Horia Velicu, Alexandru Preda, Serban Chiricescu
arXiv:2504. 07385v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) become increasingly used for question-answering (QA), relying on static, pre-annotated references for evaluation poses significant challenges in cost, scalability, and completeness.
By Sher Badshah, Ali Emami, Hassan Sajjad
arXiv:2607. 26512v1 Announce Type: new Abstract: AI agents can draft claims faster than authors can check whether the cited or retrieved evidence supports them.
By Gengyu Chen, Yongjie Yu, Weiling Wang
arXiv:2609.07075v1 Announce Type: new
Abstract: Retrieval-augmented generation (RAG) is commonly evaluated by whether the final answer is correct. That test is insufficient: an answer can match its r...
By Ramon Gonzalez, Antonio Diaz
SearchAtlas is a framework that transforms raw search trajectories of large language model (LLM) agents into structured evidential query graphs, where edges capture how evidence is propagated from queries to the final answer. The automated parsing pipeline achieves a mean edge F1 of 86.0% against human-annotated graphs and remains consistent across repeated runs. Using SearchAtlas, the authors analyze five search agents on three benchmarks, uncovering systematic differences in search scale and evidence aggregation, and revealing process failures such as fragmented answer support, unmet question constraints, and unverified parametric knowledge that correlate strongly with incorrect answers.
By Jiacheng Sang, Mengyuan Li, Sanxing Chen, Yukun Huang, Yu Feng, Bhuwan Dhingra
arXiv:2607. 09489v1 Announce Type: new Abstract: An AI system's output is not the fact or world state it appears to describe, but rather an engineered representation.
By Jade Alglave, Patrick Cousot
arXiv:2607. 17291v1 Announce Type: new Abstract: Deep research agents increasingly operate over the open web, where relevant records coexist with redundant summaries, outdated reports, and misleading documents.
By Jun Nie, Zhiqin Yang, Zhenheng Tang, Yonggang Zhang, Xiaowen Chu, Xinmei Tian, Bo Han
Large language models are increasingly deployed as agents that reason over documents rather than answer from parametric knowledge. We study archive-grounded reasoning: locating sparse evidence across a large, messy collection of workplace files, reconciling inconsistent terminology, units, and time conventions, and computing an answer.