Language models expose internal signals that predict whether an answer is correct, readable from a single forward pass of a frozen model without additional generations. Yet existing probes often commi...
The paper introduces Declarative Attention (DA), a protocol that lets language models explicitly declare which parts of their context to focus on during generation. By partitioning decoding into full-context, region-specific, and recent-output-only modes, the inference engine can skip large portions of the KV cache, dramatically reducing attended tokens. Experiments on 15 long-context tasks with off-the-shelf models show significant savings (52.0% and 31.1% reductions) with only modest accuracy drops that diminish as model size increases.
By Namgyu Ho, Huzama Ahmad, Woosung Koh, Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos
The paper introduces a method for determining whether retrieval-augmented generation (RAG) systems have sufficient, insufficient, or conflicting evidence to answer a question. By training a lightweight linear classifier on hidden activations and attention-derived features from 16 language models, the authors demonstrate that these internal signals reliably predict the adequacy of retrieved documents, outperforming prompting-based baselines and specialized RAG models. Analysis shows that middle-layer hidden states carry the most informative signals for this triage task.
By Syed Mahbubul Huq, Christopher Child, Tillman Weyde, Pranava Madhyastha
arXiv:2609.37491v1 Announce Type: cross
Abstract: Retrieval-augmented language models are expected to answer from the retrieved evidence, but in practice they often keep answering when that evidence...
By Zeyan Li, Qirong Guo, SIyuan Qiu, Hu Xu, Chun Li, Jianfeng Xu
arXiv:2607. 04223v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) reduces but does not eliminate hallucination, and existing detectors return a single answer-level score that does not indicate which sentence is unsupported, or why.
By Mohamed Aly Bouke
arXiv:2606. 27359v1 Announce Type: cross Abstract: Many decoding methods for large language models can be understood as shifting probability mass toward outputs that are more likely under the model, either locally at the token level or globally at the sequence level.
By Johannes Zenn, Jonas Geiping