arXiv AI

When Reasoning Goes Astray: Attention Dynamics of Uncontrolled Reasoning

The paper introduces RADAR, a method that models large reasoning models (LRMs) as four states and uses dynamic attention responses to detect uncontrolled reasoning in real time. RADAR identifies abnormal attention patterns that precede repetitive loops, and the authors demonstrate that realigning these patterns reduces looping while maintaining performance. The study offers a mechanistic explanation of how benign reasoning can degenerate into harmful behavior and provides actionable guidance for runtime interventions.

arXiv AI
Jun 3

Thinking Past the Answer: Evaluating Harmful Overthinking in Large Reasoning Models

arXiv:2606. 02835v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) improve performance by generating explicit intermediate reasoning traces through increased test-time compute, yet the assumption that longer reasoning is consistently beneficial remains under-examined.

By Simone Caldarella, Davide Talon, Rahaf Aljundi, Elisa Ricci, Massimiliano Mancini
arXiv AI
Sep 3

When Can Large Reasoning Models Save Thinking? Mechanistic Analysis of Behavioral Divergence in Reasoning

The paper investigates why large reasoning models (LRMs) often continue to think even when prompted to stop, a phenomenon called "Still-thinking". By examining confidence at the thinking-termination boundary, internal attention divergences, and attention allocation across prompt segments, the authors find that high perplexity and greater attention to the original question correlate with continued thinking. They propose an attention‑intervention method that suppresses explicit reasoning, which reduces inefficiency but also lowers accuracy, underscoring a trade‑off between instruction compliance, inference speed, and correctness.

By Rongzhi Zhu, Yi Liu, Jiancheng Wang, Xiangyu Liu, Zequn Sun, Yiwei Wang, Yu Deng, Zijian Zhou, Wei Hu
arXiv Machine Learning
Jun 2

Are Large Reasoning Models Interruptible?

arXiv:2510. 11713v4 Announce Type: replace-cross Abstract: Real-world applications of Large Reasoning Models (LRMs) often require reasoning about changing prompts or environments.

By Tsung-Han Wu, Mihran Miroyan, David M. Chan, Trevor Darrell, Narges Norouzi, Joseph E. Gonzalez
arXiv AI
Sep 4

</think> Doesn't Stop Reasoning: Analysis of Spurious CoT Termination

The paper investigates a training‑free early‑exit technique that inserts an end‑of‑think (EoT) token to terminate chain‑of‑thought (CoT) reasoning in large reasoning models. It finds that the injected EoT often fails to cleanly switch the model from reasoning to answering, leading to continued reasoning‑like generation—termed spurious CoT termination—whose length scales with the amount of reasoning saved. By increasing attention to the EoT token through Exit‑token Attention Biasing (EAB), the authors reduce spurious termination and shorten the answering phase across multiple models and benchmarks.

By Seunghee Koh, Sungjae Choi, Minchan Kwon, Sunghyun Baek, Junmo Kim
arXiv Machine Learning
Jun 9

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization

arXiv:2510. 13554v2 Announce Type: replace-cross Abstract: The reasoning pattern of Large language models (LLMs) remains opaque, and reinforcement learning (RL) typically applies uniform credit across an entire generation, blurring the distinction between pivotal and routine steps.

By Yang Li, Zhichen Dong, Yuhan Sun, Weixun Wang, Shaopan Xiong, Yijia Luo, Jiashun Liu, Han Lu, Jiamang Wang, Wenbo Su, Bo Zheng, Junchi Yan