arXiv Machine Learning

Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs

arXiv:2605. 09239v2 Announce Type: replace-cross Abstract: Large language models fail at counting how many times a word repeats in a list, even though they perform well on far harder reasoning tasks.

arXiv AI
Aug 25

Lexical Perturbations Disrupt LLM Reasoning: An Empirical Study of Attention Diversion

The study investigates how lexical perturbations—such as keyboard noise, character swaps, and filler insertion—affect large language models (LLMs) on reasoning benchmarks. Four open-weight instruction-tuned models and frontier models were evaluated, revealing that character-level perturbations significantly reduce accuracy, especially on multi-step reasoning tasks, while filler insertion has minimal impact. The authors attribute this asymmetry to Attention Diversion, where fragmented subword tokenization draws disproportionate attention in middle and final transformer layers; they demonstrate that both token content and attention allocation are coupled, making it difficult for inference-time repair strategies to fully recover performance.

By Jiaqian Zhu, Yang Zhang, Junhua Ding, Xiaowei Yu
arXiv AI
Sep 3

Do VLMs Read or Rewrite? On Transcription Faithfulness in Vision-Language Models

Vision‑Language Models (VLMs) are increasingly replacing traditional OCR for document understanding, but this study shows they often rewrite imperfect text into more plausible forms, a flaw that clean‑text OCR benchmarks miss. The authors created FaithC4, a multilingual perturbation benchmark of 1,455 single‑page documents with scramble, random substitution, and visually similar substitution attacks, and evaluated 15 systems across general‑purpose VLMs, OCR‑specialized VLMs, and traditional OCR pipelines. Results reveal that general‑purpose VLMs suffer up to 6.9 WER points under perturbation, OCR‑specialized VLMs 0.1–3.4 points, and traditional OCR less than 0.8 points on English; probing Qwen3‑VL‑4B shows rewriting occurs only when a perturbed word’s final‑layer representation remains close to the original, with short words (4–6 characters) rewritten up to 10% of the time. whyItMatters":"The findings highlight a critical limitation of VLMs in document transcription, underscoring the need for robust evaluation benchmarks that capture rewriting behavior beyond clean‑text accuracy."

By Gwang Gook Lee, Kenan Emir Ak, Jay Mohta, Yan Xu, Dimitrios Dimitriadis
arXiv Computation and Language
Sep 4

How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models

The paper investigates how six naturalistic and synthetic input perturbations affect decoder‑only language models at three levels: output behavior, hidden‑state geometry, and attention‑head function. Using GPT‑2 and Qwen2.5 checkpoints, the authors analyze layerwise geometry with centered kernel alignment and intrinsic dimension, and examine attention‑head responses in GPT‑2. They find that perturbation types produce distinct metric profiles that are not fully captured by output measures and vary across checkpoints, highlighting the need for multi‑level evaluation of robustness.

By Dun Li Chan, Emily Liu, Niyathi Allu, Christian Hoang
arXiv AI
Sep 25

Reasoning Instructions Can Break Answer Decoding in Vision--Language Models

The paper shows that chain‑of‑thought (CoT) instructions can distort evaluation of vision‑language models (VLMs) when a scorer reads answer‑label logits before the model generates a rationale. On ScienceQA, Qwen2.5‑VL‑7B’s accuracy falls from 80.76% to 45.48% under this CoT‑prefix scoring, and most predictions incorrectly pick the first option. Linear probes and free generation recover most of the lost accuracy, indicating that the answer information remains in the hidden states but is missed by the early readout. The authors explain the mismatch with vocabulary and layer diagnostics, noting that probability mass shifts toward continuation tokens while answer information stays linearly accessible in later layers. The effect varies across datasets and models, but the study demonstrates that CoT‑prefix scoring can misrepresent model knowledge unless the requested and scored outputs are aligned.

By Zeyan Li, Siyuan Qiu, Jianfeng Xu
Hugging Face Trending Papers
Sep 24

Reasoning Instructions Can Break Answer Decoding in Vision--Language Models

The paper shows that chain‑of‑thought (CoT) instructions can distort multiple‑choice vision‑language model evaluation when a scorer appends a reasoning cue but reads answer‑label logits before the model generates any rationale. This CoT‑prefix scoring causes significant drops in accuracy (e.g., Qwen2.5‑VL‑7B falls from 80.76% to 45.48% on ScienceQA) and leads most predictions to choose the first option. Analysis reveals that while answer information remains linearly accessible in late layers, the immediate readout is misled by probability mass shifting toward continuation tokens, and the issue varies across datasets and models.