arXiv AI

Before Reasoning Can Fail: Pre-Evidence Procedural Failures in Agentic RAG

arXiv:2608. 02011v2 Announce Type: replace Abstract: Agentic retrieval-augmented generation (RAG) systems can fail before evidence-conditioned reasoning is tested: an agent may retrieve candidate snippets but finalize without inspecting them.

arXiv AI
Aug 24

When Failures Propagate: Causal Failure Attribution in Agentic Retrieval-Augmented Generation

The paper introduces AgenticRAG-FP, an interventional benchmark designed to attribute causal failures in agentic retrieval‑augmented generation (RAG) systems. By injecting a certified fault at a specified hop and re‑executing the downstream trajectory, the benchmark evaluates whether post‑hoc diagnostics can correctly identify the fault’s location. Experiments on MuSiQue questions show that coverage‑based diagnosis performs well at hop 1 but poorly at later hops, while counterfactual probes reveal varying diagnostic success depending on propagation depth.

By Lauren Pothuru
arXiv AI
4d ago

REALHOP: Rethinking Multi-Hop Reasoning Evaluation via Behavioral Auditing

REALHOP introduces a behavioral auditing framework to assess multi‑hop reasoning by measuring the Behavioral Necessity Rate (BNR), which quantifies how often removing targeted evidence prevents correct answers. Across five benchmarks, the framework reveals a wide gap between annotated reasoning chains and actual evidence dependence, with panel‑mean BNR ranging from 16.6% to 48.9%. By re‑binding entities, factorizing relations, adding competing paths, and placing evidence at traceable locations, REALHOP raises BNR dramatically—from 27.4% to 94.4% on MuSiQue questions—while maintaining high overall accuracy and improving performance on long‑context tasks.

By Jiawen Tao, Xiaokun Yuan, Yaoming Li, Chenxu Liu, Mengzhou Wu, Tong Yang, Maxm Pan
Hugging Face Trending Papers
Sep 2

VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement

VeriPhy is an auditable physical‑verification system that evaluates generated video by compiling prompts into typed physical obligations and a statically validated execution plan before any frames are observed. During execution, it gates calls to frozen low‑level experts (e.g., segmentation, tracking, counting, depth, OCR, audio‑event detection) and returns provenance‑carrying evidence records, which are mapped to a three‑valued state (supported, contradicted, unknown) with full traceability. On a 1,500‑clip corpus of human‑annotated flaw records, VeriPhy accounts for 228 failures out of 304, outperforming a published question‑decomposition evaluator that accounts for 164, while also providing auditable evidence for each verdict.

arXiv AI
Jul 21

From Evidence to Trajectory: Abductive Reasoning Path Synthesis for Retrieval-Augmented Generation Agents Development

arXiv:2509. 23071v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) agent development is hindered by the lack of executable ground-truth agent-environment interaction trajectories.

By Muzhi Li, Jinhu Qi, Yihong Wu, Minghao Zhao, Liheng Ma, Yifan Li, Xinyu Wang, Zhenghan Tai, Zixing Song, Yingxue Zhang, Ho-fung Leung, Irwin King
arXiv AI
Sep 10

Do Web Agents Investigate Before They Decide?

The paper introduces MIRAGE, a benchmark of 750 multi‑step decision tasks designed to test autonomous web agents’ investigative abilities across Wikipedia Forensics, Shopping Admin adjudication, and Reddit Moderation. Each task contains a misleading visible context and a hidden context that holds decisive evidence, allowing performance to be broken down into Investigation, Reasoning, and Decision Accuracy, with an added Investigative Hallucination Rate. Evaluation of eight LLM agents reveals three consistent patterns: agents often reach relevant pages but fail to extract decisive evidence, procedural hints improve investigation but not decision accuracy on Wikipedia tasks, and 12.6% of trajectories include fabricated facts.

By Syed Nazmus Sakib, Nafiul Haque, Tapodhir Karmakar Taton, Shahrear Bin Amin, Shifat E. Arman
arXiv AI
Aug 11

Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills

arXiv:2608. 07885v1 Announce Type: new Abstract: Reasoning modes of language models outperform their non-reasoning counterparts on multi-step agentic tasks, but pay a 3-6x premium in output tokens on every episode -- much of it spent re-deriving procedures that are shared across episodes of the same domain.

By Agamdeep Singh, Srishti Gautam, Priyanshu Gupta, Nikita Mehrotra, Tanmay Bakshi, Sumit Gulwani