arXiv AI

Where's the Plan? Locating Latent Planning in Language Models with Lightweight Mechanistic Interventions

arXiv:2605. 07984v2 Announce Type: replace-cross Abstract: We study planning site formation in language models -- where internal representations of structurally-constrained future tokens form during the forward pass, and whether they causally drive generation.

arXiv AI
Sep 21

When Steering Fails in Latent Reasoning: A Latent-to-Language Transition Gap

The paper investigates the effectiveness of activation steering in latent chain-of-thought (CoT) reasoning compared to explicit CoT. It finds that steering continuous latent thoughts yields weaker impacts on language generation, even when hidden representations are shifted similarly. The authors propose a latent-to-language transition gap, supported by evidence of abrupt output distribution changes at the transition boundary and weaker bidirectional control in latent CoT.

By Gaoxiang Huang, Lei Qi
arXiv Computation and Language
Aug 25

Mechanistic Interpretability of Chain-of-Thought Reasoning via Sequential Activation Patching

The paper introduces a sequential activation patching framework to study how Chain-of-Thought (CoT) prompting influences large language models over multiple generated tokens. By tracking CoT-conditioned attention-head activations across token positions and aggregating them with Part-of-Speech guidance, the authors identify distributed head sets that jointly contribute to answer generation. Targeted zero-ablation experiments confirm that these heads are functionally important, affecting mechanisms such as reasoning-trajectory maintenance, answer anchoring, exemplar-target separation, and numerical generation.

By Murat Dura, Serkan \"Ozt\"urk, Selma Tekir
arXiv AI
Jul 22

Fluid Reasoning Representations

arXiv:2602. 04843v2 Announce Type: replace Abstract: Frontier large language models increasingly solve complex tasks involving abstract concepts through extended test-time thinking.

By Dmitrii Kharlapenko, Terry Jingchen Zhang, Arth Singh, Alessandro Stolfo, Arthur Conmy, Mrinmaya Sachan, Zhijing Jin
arXiv AI
Sep 10

A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM

A*-Thought-V2 is a framework that models Chain-of-Thought reasoning as a geometric trajectory in a 3D PCA space, using explicit-implicit latent tokens to compress steps that deviate from the main question-to-solution direction. The method measures alignment angles to decide which steps remain text and which become latent, and introduces stepwise embedding forcing and label forcing to train the architecture. Experiments on Qwen models show up to 2.6% accuracy gains, halved response length, and significant reductions in computation and training time.

By Xiaoang Xu, Siyuan Liu, Shuo Wang, Junlan Feng, Fanyu Meng, Zhu Zhang, Jixun Wang, Xiaorong Wang, Zihan Zhou, Xin Li, Chaojun Xiao, Yiming Zhang, Huijia Wu, Liuyu Xiang, Peipei Li, Zhaofeng He
arXiv Machine Learning
Jun 5

Temporal Preference Concepts and their Functions in a Large Language Model

arXiv:2606. 05194v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly being deployed to make decisions that require trading off near-term gains against long-term consequences, yet little is known about how they internally represent or resolve these tradeoffs.

By Ian Rios-Sialer, Shantanu Darveshi, Shuai Jiang, Avigya Paudel, Anastasiia Pronina, Ipshita Bandyopadhyay, Justin Shenk
arXiv AI
Sep 15

Towards a Mechanistic Understanding of Propositional Logical Reasoning in Large Language Models

The paper investigates how large language models perform propositional logical reasoning by conducting a causal mechanistic analysis on the PropLogic-MI benchmark. It identifies four interlocking mechanisms—Staged Computation, Information Transmission, Fact Retrospection, and Specialized Attention Heads—that organize the reasoning process across layers. The study demonstrates that these mechanisms recur across different model families, rule categories, and reasoning hops, indicating a structured, layer‑organized internal process for propositional reasoning.

By Danchun Chen, Qiyao Yan, Chenpeng Wang, Liangming Pan