arXiv Machine Learning By Wenzhang Du

What Would Fix This RAG Failure? Auditing Counterfactual Response with Paired Evidence Interventions

Read the original on arXiv Machine Learning →

arXiv:2608. 08944v1 Announce Type: cross Abstract: A failed retrieval-augmented generation (RAG) answer can be consistent with several unseen responses to evidence repair.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 18

Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression

The paper investigates how memory systems can answer a current query correctly yet fail to retain distinctions needed for later updates. Using a paired‑history audit, the authors evaluate 24 history pairs across six synthetic mechanisms and two model backends, achieving perfect reveal accuracy on DeepSeek and high accuracy on GLM. Record‑level audits reveal specific failures in structured reveal memories and frontier late‑reference adequacy, and the authors test a label‑equivariant repair that only partially restores correctness.

By Guangzhe Zhang
arXiv AI
Aug 24

When Failures Propagate: Causal Failure Attribution in Agentic Retrieval-Augmented Generation

The paper introduces AgenticRAG-FP, an interventional benchmark designed to attribute causal failures in agentic retrieval‑augmented generation (RAG) systems. By injecting a certified fault at a specified hop and re‑executing the downstream trajectory, the benchmark evaluates whether post‑hoc diagnostics can correctly identify the fault’s location. Experiments on MuSiQue questions show that coverage‑based diagnosis performs well at hop 1 but poorly at later hops, while counterfactual probes reveal varying diagnostic success depending on propagation depth.

By Lauren Pothuru