arXiv AI

The Model Knows Your Project, Not You: Measuring Recognition in LLMs with NameRank

arXiv:2607. 12520v1 Announce Type: new Abstract: What a frontier model recalls about a person or tool from its own weights -- before any retrieval step -- often shapes the first description a human sees, making that parametric corpus presence a measurement problem.

arXiv AI
2d ago

On-Device Named-Entity Recognition: A Deployability Study of Accuracy, Cost, Reliability, and Confidence

The paper evaluates nine on‑device named‑entity recognition models ranging from classical taggers to large language models, measuring not only accuracy but also latency and output validity. Using a silver‑gold benchmark derived from an LLM judge panel and a human‑validated corpus, the study shows that encoder‑based models achieve comparable accuracy to a 4 B instruct LLM while being much smaller, faster, and producing no malformed output. Confidence calibration of GLiNER is analyzed, revealing over‑confidence but improved reliability after temperature scaling and thresholding.

By Vinay Kumar Chaganti
arXiv Machine Learning
Sep 11

Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models

The paper investigates whether language models can identify sentences from their training data by using exact duplication counts from publicly released corpora for two model families, OLMo‑2 and Pythia. It finds that for typical duplication levels, models show only a weak trace of exposure, with a rank correlation near –0.08, and that strong signals only appear when a sentence appears roughly a thousand times, at which point fame rather than memory dominates. The study also demonstrates that common membership tests can be misleading, as changing a single word does not alter the model’s preference, and that controlling for register can significantly improve detector performance.

By Arman Nik Khah
arXiv AI
Sep 10

No\=esis: Deterministic-First Retrieval with Two-Tier Context Hydration for Factuality-Critical Queries on Small Local Models

Noesis is a retrieval-augmented architecture designed for small local language models that prioritizes deterministic judgments before generation. It employs a fact layer, positional addressing, provenance scoping, and two-tier context hydration to ensure factual integrity, achieving parity with larger models on exact value accuracy and eliminating confabulated numbers. The system delivers single-generation queries with traceable, source‑linked outputs, improving performance in regulated domains.

By Nicola Cogotti
arXiv AI
Aug 26

Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes

The paper audits a 366‑day autobiographical book generated by a large language model (LLM) against an independent verification corpus. Using a four‑level rubric, 354 of the 366 days (96.7%) failed verification, with only 12 days containing corroborated scenes and 19 days containing actively contradicted claims. Regenerating the same days with current models yielded 100% verification failure, while grounding the generation in the subject’s own corpus improved the rate to 83.3% but still left substantial residual failure.

By Heather Renze
arXiv AI
Sep 7

Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model Pipelines

The paper investigates whether the object selected in a grounded language‑model pipeline actually reaches the reader, a failure that can break the handoff between stages. By auditing 600 HybridQA questions across three selector families, the authors find that exact key lookup and title matching recover the selected object in all 1,463 resolvable records, but body‑only BM25 omits it in 26.6% of cases at cutoff five, while hybrid retrieval with reranking omits it only 1.0%. The study also shows that misalignment between selected and retrieved objects can reduce exact match scores by up to 31 points, and introduces the Returned‑Object Profile (ROP) as a tool for reproducible auditing.

By Siddharth Vohra, Runmin Jiang, Xiaomo Li, Min Xu
arXiv Computation and Language
Sep 22

Re:CAP - Auditing Retrieval Coverage in Production RAG Pipelines

Re:CAP is a reference‑free audit loop for retrieval‑augmented generation (RAG) pipelines that probes for missing documents instead of enumerating all relevant ones. It identifies covered topics, generates probing questions, retrieves candidate documents, and uses an LLM judge to keep only those that add new information. On several benchmarks, Re:CAP recovers a significant portion of gold documents that flat BM25 or hybrid retrieval misses, and human evaluation shows most of these documents add new information.

By Aviral Joshi, Hanoz Bhathena, Max Nelson, Saket Sharma