arXiv AI

Universal Pathologies, Conditional Consequences: A Triple-Robustness Analysis of RAG for Multi-Hop Traceability

arXiv:2608. 05153v1 Announce Type: cross Abstract: GraphRAG underperforms vector RAG on citation precision in many reports, but where and why have remained corpus-bound.

arXiv AI
Jul 1

RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora

arXiv:2604. 19047v2 Announce Type: replace-cross Abstract: Existing QA benchmarks typically assume distinct documents with minimal overlap, yet real-world retrieval-augmented generation (RAG) systems operate on corpora such as financial reports, legal codes, and patents, where information is highly redundant and documents exhibit strong inter-document similarity.

By Hanjun Cho, Jay-Yoon Lee
arXiv Computation and Language
Sep 22

Re:CAP - Auditing Retrieval Coverage in Production RAG Pipelines

Re:CAP is a reference‑free audit loop for retrieval‑augmented generation (RAG) pipelines that probes for missing documents instead of enumerating all relevant ones. It identifies covered topics, generates probing questions, retrieves candidate documents, and uses an LLM judge to keep only those that add new information. On several benchmarks, Re:CAP recovers a significant portion of gold documents that flat BM25 or hybrid retrieval misses, and human evaluation shows most of these documents add new information.

By Aviral Joshi, Hanoz Bhathena, Max Nelson, Saket Sharma
arXiv Machine Learning
Sep 21

Predictable Failure in Multi-Hop Retrieval: Score-Distributional Confidence Scoring and Abstention

The paper shows that multi‑hop retrieval failures cluster in predictable subpopulations and formalizes this with two theoretical results: (1) confident‑failure reduction is possible only when retrieval features carry mutual information about success, and (2) no single ANN score feature dominates across all failure regimes. Building on these insights, the authors introduce RegimeAbstain, which computes a Retrieval Confidence Score (RCS) from up to nine query‑ANN structural features and uses it to calibrate an abstention policy. Across three benchmarks and two retrieval architectures, RCS achieves the best or co‑best AUC‑AC and significantly reduces the Confident‑Wrong‑Answer Rate, demonstrating its effectiveness and domain‑agnostic applicability.

By Andre Bacellar
arXiv Computation and Language
Sep 23

ABAI at COLIEE 2026 Task 1: Multi-Stage Retrieval with GraphRAG-Enhanced Meta-Learning, and a Post-Hoc Study of the Cross-Validation-to-Test Gap

The paper reports the ABAI submission to COLIEE 2026 Task 1, a case law retrieval challenge that suppresses cited passages, and details a four‑stage retrieval pipeline: multi‑view BM25 with reciprocal rank fusion, neural reranking, graph‑based features via a graph attention network, and a LightGBM meta‑learner over 34 features. The best run achieved an F1 score of 0.177 on the official test set, compared to a cross‑validated 0.311, and the authors attribute the gap to a recall ceiling, temporal distribution shift, and threshold miscalibration. A controlled post‑hoc study examined the impact of threshold transfer, decision quality across time, and query similarity, and identified specific remedies—such as BM25 length‑normalisation tuning, event‑triple views, and dense fusion—that improved recall, while other interventions had no effect.

By Minhan Cho, Soyoung Park, Daejin Choi, Jinyoung Han
arXiv Computation and Language
4d ago

Pair Difficulty Matters: Rethinking Pairwise LLM-as-a-Judge Evaluation and Consistency

Large Language Model judges are commonly used to rank texts via pairwise comparison, with reliability traditionally measured by position bias, transitivity, and pairwise agreement. This paper argues that these proxies are misleading because they are dominated by close‑rank‑gap pairs, which contribute little to the overall ranking, while far‑gap pairs carry the true ranking signal. Experiments on simulations and human‑rated corpora show weak correlation between the proxies and actual ranking accuracy, suggesting judges should be evaluated using rank‑gap‑conditional metrics against human rankings.

By Bruno Brocai, Maria Becker
arXiv AI
Sep 1

post-graph-rag: A PostgreSQL-Native Bi-Temporal Graph RAG Engine with Temporal Grounding at Synthesis

post-graph-rag is an open‑source PostgreSQL‑native engine that unifies chunks, embeddings, a canonical entity graph, and community summaries in a single database, using pgvector for search and edge tables for traversal. It validates extraction output—rejecting vague predicates, normalising predicates, resolving entities to unique vertices, and flagging negations—before writing, and employs a bi‑temporal layer to record when a relation held and when the system believed it, superseding incompatible earlier assertions. In benchmarks against LightRAG, it builds denser, more queryable graphs and achieves higher scores on LongMemEval, largely due to its temporal grounding in prompts.

By Chandan Rajah
arXiv AI
Sep 24

Meet, Compare, or Abstain: LatWeave for Deterministic Multi-Hop Question Answering on Knowledge Lattices

LatWeave is a deterministic multi‑hop question‑answering framework that structures knowledge into a multidimensional lattice and reduces QA to three operators—meet, compare, and abstain—while limiting LLM use to extraction and planning. It achieves near‑lossless performance on complete knowledge benchmarks (e.g., MetaQA) and strong results on templated multi‑hop datasets (e.g., 2WikiMultihopQA), while transparently handling incomplete knowledge through abstention. The approach offers reproducible, auditable answer paths with no performance penalty within its operating envelope.

By Yuze Ren, Shaoheng Fan, Tao Wang, Yabo Yan, Han Han