arXiv Computation and Language

Stale-Document Poisoning: When Outdated Retrieval Overrides Correct Model Answers

The paper introduces "stale‑document poisoning," a temporal alignment failure where outdated retrieval evidence causes retrieval‑augmented generation models to produce incorrect answers even when the model could answer correctly without retrieval. A benchmark of 317 verified knowledge reversals across medicine, law, software, and platform policy shows that outdated evidence flips 30–37% of Llama and Qwen answers, rising to 66–75% when models are explicitly instructed to trust the document. The study demonstrates that providing explicit validity information and a recency‑aware re‑ranker can substantially reduce poisoning, highlighting the need for models to assess whether retrieved evidence remains applicable.

arXiv Computation and Language
4d ago

Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification

The paper investigates why retrieval‑based open‑ended evaluation fails in medical fact verification. By creating two detailed taxonomies—one for retrieval‑stage errors across five quality dimensions and another for verifier‑reasoning errors across six steps—the authors automatically label evidence quality and reasoning errors using an LLM‑as‑Judge pipeline. Their large‑scale stress tests across multiple retrieval methods and verifier models show that increasing model size, reasoning effort, source breadth, or medical fine‑tuning does not eliminate these failure modes, indicating fundamental limits of the retrieve‑then‑verify paradigm in open‑ended medical contexts.

By Heyuan Huang, Jirui Dai, Alexandra DeLucia, Sonal Joshi, Mahsa Yarmohammadi, Jie Gao, Bernal Jim\'enez Guti\'errez, Mark Dredze
arXiv AI
Sep 18

Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression

The paper investigates how memory systems can answer a current query correctly yet fail to retain distinctions needed for later updates. Using a paired‑history audit, the authors evaluate 24 history pairs across six synthetic mechanisms and two model backends, achieving perfect reveal accuracy on DeepSeek and high accuracy on GLM. Record‑level audits reveal specific failures in structured reveal memories and frontier late‑reference adequacy, and the authors test a label‑equivariant repair that only partially restores correctness.

By Guangzhe Zhang
arXiv AI
Aug 11

Time Present and Time Past: Benchmarking Large Language Models on Temporally Evolving Document Understanding

arXiv:2608. 08512v1 Announce Type: new Abstract: Evolving documents, such as laws, tax codes, and software documentation, are amended, replaced, and sometimes reverted over time, so a question has different correct answers at different dates.

By Mahbub E Sobhani, Md. Faiyaz Abdullah Sayeedi, Fahmid Hasan Chowdhury, Md Adnan Arefeen, Farig Sadeque, Md. Faizul Bari, Swakkhar Shatabda
arXiv Machine Learning
Sep 25

Return or Revise? Learning When Revision Helps Retrieval-Augmented QA

The paper investigates when it is better to return an existing draft answer or revise it using retrieved evidence in retrieval‑augmented QA systems. By grading both the draft and its candidate revision with the same correctness judge, the authors define a paired effect called recoverability and train policies to predict it before revision. Experiments on 25,870 open‑domain questions show that a recoverability‑based scorer outperforms a draft‑correctness scorer across multiple Llama setups, improving accuracy–revision trade‑offs and closing a significant portion of the oracle gap, though it still applies harmful revisions in a substantial fraction of cases.

By Nicholas Kashani Motlagh, Tim Anderson, Jeremy Gwinnup, Grant Erdmann
arXiv AI
2d ago

When Tools Silently Lie: Evaluating and Mitigating Blind Compliance in Tool-Augmented Data Agents

The paper introduces ToxicBench, a benchmark designed to evaluate how tool‑augmented data agents handle incorrect tool outputs. By pairing clean and poisoned observations across numerical, label, schema, and retrieval errors, the authors assess both the checking process and the final answer adoption. In a 118‑task GPT evaluation, poisoning reduces task success by 26–39 percentage points, revealing that repeated poisoning leads to wrong-answer adoption even after checking, while ordinary retries help only under one‑shot poisoning. Human annotations on 200 trajectories confirm the scoring system’s reliability, showing 96% agreement with task success and supporting the benefits of retries and audit‑based adoption.

By Zifu Tao, Changqing Yin