The paper shows that chain‑of‑thought (CoT) instructions can distort multiple‑choice vision‑language model evaluation when a scorer appends a reasoning cue but reads answer‑label logits before the model generates any rationale. This CoT‑prefix scoring causes significant drops in accuracy (e.g., Qwen2.5‑VL‑7B falls from 80.76% to 45.48% on ScienceQA) and leads most predictions to choose the first option. Analysis reveals that while answer information remains linearly accessible in late layers, the immediate readout is misled by probability mass shifting toward continuation tokens, and the issue varies across datasets and models.
The paper shows that chain‑of‑thought (CoT) instructions can distort evaluation of vision‑language models (VLMs) when a scorer reads answer‑label logits before the model generates a rationale. On ScienceQA, Qwen2.5‑VL‑7B’s accuracy falls from 80.76% to 45.48% under this CoT‑prefix scoring, and most predictions incorrectly pick the first option. Linear probes and free generation recover most of the lost accuracy, indicating that the answer information remains in the hidden states but is missed by the early readout. The authors explain the mismatch with vocabulary and layer diagnostics, noting that probability mass shifts toward continuation tokens while answer information stays linearly accessible in later layers. The effect varies across datasets and models, but the study demonstrates that CoT‑prefix scoring can misrepresent model knowledge unless the requested and scored outputs are aligned.
By Zeyan Li, Siyuan Qiu, Jianfeng Xu
arXiv:2608. 11694v1 Announce Type: cross Abstract: A benchmark score comes from a single phrasing of each problem.
By Shailja Thakur, Sungeun An, Chad DeLuca, Hima Patel
arXiv:2609.21383v1 Announce Type: cross
Abstract: Depth-recurrent language models refine predictions through repeated latent updates. Why can intermediate answers agree with the endpoint while their...
By Xinyue Luo, Fei Yu
The paper investigates where and how large language models encode signals that indicate answer correctness. By examining hidden states, token probabilities, residual-stream features, attention, and their combinations, the authors find that correctness signals are concentrated in the answer span and that different signal families complement each other. Fusing these signals improves robustness, especially under distribution shifts, and can be used to control retrieval in downstream tasks.
By Yuansen Liu, Yixuan Tang, Anthony Kum Hoe Tung
arXiv:2605. 26795v2 Announce Type: replace Abstract: Chain-of-thought (CoT) prompting enhances large language model performance, yet what drives these gains remains unclear.
By Xiang Wang, Wei Wei