arXiv AI
Jul 1

RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora

arXiv:2604. 19047v2 Announce Type: replace-cross Abstract: Existing QA benchmarks typically assume distinct documents with minimal overlap, yet real-world retrieval-augmented generation (RAG) systems operate on corpora such as financial reports, legal codes, and patents, where information is highly redundant and documents exhibit strong inter-document similarity.

By Hanjun Cho, Jay-Yoon Lee
arXiv AI
Sep 10

No\=esis: Deterministic-First Retrieval with Two-Tier Context Hydration for Factuality-Critical Queries on Small Local Models

Noesis is a retrieval-augmented architecture designed for small local language models that prioritizes deterministic judgments before generation. It employs a fact layer, positional addressing, provenance scoping, and two-tier context hydration to ensure factual integrity, achieving parity with larger models on exact value accuracy and eliminating confabulated numbers. The system delivers single-generation queries with traceable, source‑linked outputs, improving performance in regulated domains.

By Nicola Cogotti
arXiv Computation and Language
Sep 1

Annotated Surrogate Retrieval for Polish Statutory Law

The paper introduces three retrieval methods for Polish statutory law that use language‑model annotations attached to articles as surrogates. The methods—ASCR, ASCR‑H, and DTF—vary in cost and quality, with ASCR‑H achieving the highest rank‑one accuracy on bar exam questions, while DTF offers competitive performance with lower latency and cost. Extensive evaluation against 14 baselines on 300 exam questions demonstrates significant improvements in head‑rank accuracy and discusses limitations such as coverage asymmetry and negative results for lemmatisation, pseudo‑relevance feedback, and query rewriting.

By Orkun Yi\u{g}it Cengiz