arXiv AI

Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!

arXiv:2504. 09762v4 Announce Type: replace Abstract: Intermediate token generation (ITG), where a model produces output before the solution, has become a standard method to improve the performance of language models on reasoning tasks.

arXiv AI
Sep 4

</think> Doesn't Stop Reasoning: Analysis of Spurious CoT Termination

The paper investigates a training‑free early‑exit technique that inserts an end‑of‑think (EoT) token to terminate chain‑of‑thought (CoT) reasoning in large reasoning models. It finds that the injected EoT often fails to cleanly switch the model from reasoning to answering, leading to continued reasoning‑like generation—termed spurious CoT termination—whose length scales with the amount of reasoning saved. By increasing attention to the EoT token through Exit‑token Attention Biasing (EAB), the authors reduce spurious termination and shorten the answering phase across multiple models and benchmarks.

By Seunghee Koh, Sungjae Choi, Minchan Kwon, Sunghyun Baek, Junmo Kim
arXiv Computation and Language
Sep 7

Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs

The paper investigates how large language models (LLMs) organize reasoning operations—such as problem formulation, goal decomposition, and deduction—within their hidden representation spaces. It shows that these operations are separable in held‑out representations, with peak separability in middle layers, and that token‑wise alignment of operations becomes more distributed across spans as layers deepen. Attention‑masking experiments reveal that representations aligned to operations at chunk onsets depend on prior reasoning context, indicating a geometric correspondence between linguistic reasoning expressions and internal model structure.

By Seogyeong Jeong, Jaehui Hwang, Dongyoon Han, Geonmo Gu, Alice Oh, Taekyung Kim
arXiv AI
Sep 15

Thought without systematicity? Evaluating reasoning models on rule induction tasks

The paper investigates whether current reasoning models exhibit systematicity—the idea that understanding one concept should extend to closely related variations—by extending rule induction tasks from cognitive science. Using task isomorphisms like recombination and substitution, the authors generate structurally equivalent task variants and test models on them. Results show that while models can solve the original tasks, they frequently fail on these equivalent variants, indicating a lack of systematicity in their reasoning abilities.

By Simon Schug, Brenden M. Lake
arXiv AI
Jun 9

MixReasoning: Switching Modes to Think

arXiv:2510. 06052v2 Announce Type: replace Abstract: Reasoning models enhance performance by tackling problems in a step-by-step manner, decomposing them into sub-problems and exploring long chains of thought before producing an answer.

By Haiquan Lu, Gongfan Fang, Xinyin Ma, Qi Li, Xinchao Wang