arXiv Computation and Language

The Copy Ceiling: An Input-Exposure Control for Ontology-Grounded Generation over Curated Corpora

arXiv AI
Sep 7

Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model Pipelines

The paper investigates whether the object selected in a grounded language‑model pipeline actually reaches the reader, a failure that can break the handoff between stages. By auditing 600 HybridQA questions across three selector families, the authors find that exact key lookup and title matching recover the selected object in all 1,463 resolvable records, but body‑only BM25 omits it in 26.6% of cases at cutoff five, while hybrid retrieval with reranking omits it only 1.0%. The study also shows that misalignment between selected and retrieved objects can reduce exact match scores by up to 31 points, and introduces the Returned‑Object Profile (ROP) as a tool for reproducible auditing.

By Siddharth Vohra, Runmin Jiang, Xiaomo Li, Min Xu
arXiv Computation and Language
Sep 22

Re:CAP - Auditing Retrieval Coverage in Production RAG Pipelines

Re:CAP is a reference‑free audit loop for retrieval‑augmented generation (RAG) pipelines that probes for missing documents instead of enumerating all relevant ones. It identifies covered topics, generates probing questions, retrieves candidate documents, and uses an LLM judge to keep only those that add new information. On several benchmarks, Re:CAP recovers a significant portion of gold documents that flat BM25 or hybrid retrieval misses, and human evaluation shows most of these documents add new information.

By Aviral Joshi, Hanoz Bhathena, Max Nelson, Saket Sharma
arXiv Computation and Language
Sep 25

Measuring Brand and Source Discovery under Repeated LLM Queries: A Finite-Sample Audit

The study audits large language model (LLM) outputs by measuring how well repeated queries recover a collected set of responses versus the full set of possible outputs. Using sample-based rarefaction on 4,500 responses from 50 buying questions across six configurations, the authors find historical-dictionary median recovery rates between 92.6% and 95.2%, which drop to 89.5%–94.7% after re‑adjudicating all candidate strings. Additional analyses with Gemini 3.1 Pro annotations and matched roster data confirm that recovery percentages vary with extraction methods, question selection, and the finite reference collection, underscoring the need for explicit measurement definitions and sensitivity analyses in LLM audits.

By Dmitrij \.Zatuchin
arXiv Computation and Language
Aug 27

Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence

The study investigates how prior scores influence large language model (LLM) judgments in the LLM-as-a-Judge paradigm. By testing three prompt conditions—no metadata, revision framing, and anchored metadata containing prior scores—the authors find that prior scores systematically bias evaluations, shifting ratings toward those scores across 192,000 attempts. The bias also affects categorical decisions, blocking 48% of error corrections and flipping 10.18% of correct judgments, and is not mitigated by Chain-of-Thought or a warning, underscoring the need for careful context engineering.

By Ante Kapetanovic, Kemal Altwlkany, Andro Mercep, Tomislav Duricic, Emanuel Lacic
arXiv Machine Learning
Sep 11

Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models

The paper investigates whether language models can identify sentences from their training data by using exact duplication counts from publicly released corpora for two model families, OLMo‑2 and Pythia. It finds that for typical duplication levels, models show only a weak trace of exposure, with a rank correlation near –0.08, and that strong signals only appear when a sentence appears roughly a thousand times, at which point fame rather than memory dominates. The study also demonstrates that common membership tests can be misleading, as changing a single word does not alter the model’s preference, and that controlling for register can significantly improve detector performance.

By Arman Nik Khah