arXiv Machine Learning

When are likely answers right? On Sequence Probability and Correctness in LLMs

arXiv:2606. 27359v1 Announce Type: cross Abstract: Many decoding methods for large language models can be understood as shifting probability mass toward outputs that are more likely under the model, either locally at the token level or globally at the sequence level.

arXiv AI
4d ago

Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs

The paper introduces Divergent Token Confidence (DTC), a method that estimates large language model confidence by counting tokens where two models strongly disagree during decoding. DTC uses Jensen-Shannon divergence between next-token distributions along the same reasoning trajectory and shows a near-negative correlation with answer accuracy. Experiments on multiple model families and six mathematical benchmarks demonstrate that DTC improves calibration over traditional probability-based and verbalized baselines, achieving lower expected calibration errors in both white-box and black-box settings.

By Feiyang Li, Shengjing Liu, Qi Zhan, Sijie Cheng, Weiqing Wang, Hongwen Chen, Yuxuan Yang, Wen Wang, Yile Wang, Hui Huang
arXiv Computation and Language
Sep 4

Breaking the Likelihood Trap: Variance-Calibrated Modulation for Large Language Model Decoding

The paper introduces Variance‑Calibrated Modulation (VCM), a training‑free pre‑decoding technique that reshapes language model probability distributions before truncation. VCM uses two dynamic mechanisms: a Contextual Searchlight via PMI to suppress stopwords and highlight context‑relevant tokens, and an Adaptive Self‑Debiasing that applies scale‑invariant penalization based on real‑time logit standard deviation. Experiments on open‑ended generation, factual QA, and mathematical reasoning show that VCM consistently reduces the likelihood trap, improving diversity, coherence, and reasoning accuracy with minimal computational cost.

By Yuanhao Ding, Meimingwei Li, Esteban Garces Arias, Matthias A{\ss}enmacher, Christian Heumann, Chongsheng Zhang
Hugging Face Trending Papers
Sep 24

Reasoning Instructions Can Break Answer Decoding in Vision--Language Models

The paper shows that chain‑of‑thought (CoT) instructions can distort multiple‑choice vision‑language model evaluation when a scorer appends a reasoning cue but reads answer‑label logits before the model generates any rationale. This CoT‑prefix scoring causes significant drops in accuracy (e.g., Qwen2.5‑VL‑7B falls from 80.76% to 45.48% on ScienceQA) and leads most predictions to choose the first option. Analysis reveals that while answer information remains linearly accessible in late layers, the immediate readout is misled by probability mass shifting toward continuation tokens, and the issue varies across datasets and models.

arXiv AI
Sep 25

Reasoning Instructions Can Break Answer Decoding in Vision--Language Models

The paper shows that chain‑of‑thought (CoT) instructions can distort evaluation of vision‑language models (VLMs) when a scorer reads answer‑label logits before the model generates a rationale. On ScienceQA, Qwen2.5‑VL‑7B’s accuracy falls from 80.76% to 45.48% under this CoT‑prefix scoring, and most predictions incorrectly pick the first option. Linear probes and free generation recover most of the lost accuracy, indicating that the answer information remains in the hidden states but is missed by the early readout. The authors explain the mismatch with vocabulary and layer diagnostics, noting that probability mass shifts toward continuation tokens while answer information stays linearly accessible in later layers. The effect varies across datasets and models, but the study demonstrates that CoT‑prefix scoring can misrepresent model knowledge unless the requested and scored outputs are aligned.

By Zeyan Li, Siyuan Qiu, Jianfeng Xu
arXiv AI
4d ago

Locating Answer-Correctness Signals in Frozen Large Language Models

The paper investigates where and how large language models encode signals that indicate answer correctness. By examining hidden states, token probabilities, residual-stream features, attention, and their combinations, the authors find that correctness signals are concentrated in the answer span and that different signal families complement each other. Fusing these signals improves robustness, especially under distribution shifts, and can be used to control retrieval in downstream tasks.

By Yuansen Liu, Yixuan Tang, Anthony Kum Hoe Tung