arXiv:2609.24885v1 Announce Type: new
Abstract: When a language model answers from a curated corpus via graph-based retrieval, a large grounding uplift does not establish reasoning over the retrieved...
By John J. O'Hare
The paper introduces a joint fact‑verification score that evaluates both answers and the evidence submitted with them. On the FEVEROUS dataset, replacing the DCUF evidence with UnifEE evidence improves the strict score by about 9.6 percentage points, while answer accuracy rises only 1.96 points. The study also shows that increasing context length for large language models yields modest evidence‑gain improvements, and that detailed answer‑evidence analyses uncover patterns missed by aggregate metrics.
By Han Chen, Yingrui Li
The paper introduces a penalty‑aware evaluation framework for Retrieval‑Augmented Generation (RAG) systems that uses asymmetric scoring, knowledge‑gap canaries, and a failure‑attribution pipeline. Applying this framework to three commercial RAG products and a baseline on SimpleQA‑Verified, the authors find that while overall accuracy is similar across systems, canary violation rates vary dramatically, showing that systems differ more in when they answer than in what they answer. The study demonstrates that penalty‑aware scoring can reorder system rankings and is robust across different penalty settings.
By Alden Do Rosario, Hussein Younes, Felipe Pires
arXiv:2609.38021v1 Announce Type: cross
Abstract: We evaluate an auditable long-term memory system on LongMemEval-S. Its retrieval chain uses hybrid candidate retrieval, cross-encoder reranking, cove...
By Christopher J. Chanhnourack
arXiv:2605. 03534v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) grounds answers in retrieved passages, yet relevance does not guarantee sufficiency: a topical passage may still fail to justify the answer.
By Jingxi Qiu, Zeyu Han, Cheng Huang
arXiv:2607. 23804v1 Announce Type: cross Abstract: Context attribution methods for large language models (LLMs) identify which input context contributes to the model response.
By Quoc-Huy Trinh, Lin Zhu, Sebastian Szyller
arXiv:2608. 11922v1 Announce Type: cross Abstract: Predictive-distribution entropy makes a strong selection rule in retrieval-augmented question answering: across five QA benchmarks, keeping the candidate answer that a frozen respondent LLM produces with the lowest answer-token entropy lifts mean answer $F_1$ from 0.
By Po-Jen Ko, Che-Cheng Wu, Hung-Chun Hsu, Li-Yang Chang, Chuan-Ju Wang
arXiv:2609. 18960v1 Announce Type: new Abstract: Quality-aware synthetic-data selection rests on a proxy: examples that an LLM judge rates as good should also help a downstream model learn.
By Son Ha Xuan, Phat T. Tran-Truong, Xuan-Bach Le
The paper introduces the concept of LLM‑specific utility, defining it as the performance gain a target large language model (LLM) achieves when provided with a passage compared to answering without evidence. A benchmark of utilitarian passages is built for four LLMs (Qwen3‑8B/14B/32B and Llama 3.1‑8B) across three QA datasets, revealing that each model benefits most from its own tailored evidence and that evidence optimized for other models is consistently suboptimal. The authors also create SpecUBench, a benchmark for LLM‑specific utility judgment, and show that current utility‑aware retrieval methods largely capture model‑agnostic usefulness, struggling to estimate LLM‑specific utility.
"whyItMatters":"The study demonstrates that retrieval‑augmented generation must consider model‑specific evidence selection to truly improve LLM performance, highlighting a gap in existing utility‑aware methods."
By Hengran Zhang, Keping Bi, Jiafeng Guo, Jiaming Zhang, Shuaiqiang Wang, Dawei Yin, Xueqi Cheng
The study evaluates whether aggregated brand recommendation profiles can identify the language model that generated them. Using 6,475 responses from five deployed endpoints, a character‑n‑gram classifier accurately attributes single responses to the correct system (97.84% accuracy). However, when responses are aggregated into domain‑condition units, the classifier’s performance drops to 66.53%, and a forest model misclassifies all gift‑domain units, indicating that aggregated brand behaviour does not reliably reveal the underlying system.
By Dmitrij \.Zatuchin
arXiv:2609.14245v1 Announce Type: cross
Abstract: Context compression reduces generator input in retrieval-augmented generation, but answer quality alone does not characterize citation attribution. W...
By Deepanshu Mody
The paper evaluates nine on‑device named‑entity recognition models ranging from classical taggers to large language models, measuring not only accuracy but also latency and output validity. Using a silver‑gold benchmark derived from an LLM judge panel and a human‑validated corpus, the study shows that encoder‑based models achieve comparable accuracy to a 4 B instruct LLM while being much smaller, faster, and producing no malformed output. Confidence calibration of GLiNER is analyzed, revealing over‑confidence but improved reliability after temperature scaling and thresholding.
By Vinay Kumar Chaganti