arXiv:2608. 12334v1 Announce Type: cross Abstract: Despite the impressive multilingual capabilities of Large Language Models, the latent dynamics dictating language selection remain poorly understood.
By Arnav Srivastav
The paper investigates the effectiveness of activation steering in latent chain-of-thought (CoT) reasoning compared to explicit CoT. It finds that steering continuous latent thoughts yields weaker impacts on language generation, even when hidden representations are shifted similarly. The authors propose a latent-to-language transition gap, supported by evidence of abrupt output distribution changes at the transition boundary and weaker bidirectional control in latent CoT.
By Gaoxiang Huang, Lei Qi
The paper introduces a sequential activation patching framework to study how Chain-of-Thought (CoT) prompting influences large language models over multiple generated tokens. By tracking CoT-conditioned attention-head activations across token positions and aggregating them with Part-of-Speech guidance, the authors identify distributed head sets that jointly contribute to answer generation. Targeted zero-ablation experiments confirm that these heads are functionally important, affecting mechanisms such as reasoning-trajectory maintenance, answer anchoring, exemplar-target separation, and numerical generation.
By Murat Dura, Serkan \"Ozt\"urk, Selma Tekir
arXiv:2606. 01810v1 Announce Type: new Abstract: Current benchmarks for embodied vision-language planning often favor linguistic next-token prediction over physically grounded next-state reasoning.
By Zheng Lu, Mingqi Gao, Qinlei Xie, Wanqi Zhong, Hanwen Cui, Heng Cao, Zirui Song, Yifan Yang, Chong Luo, Bei Liu, Yiming Li
arXiv:2602. 04843v2 Announce Type: replace Abstract: Frontier large language models increasingly solve complex tasks involving abstract concepts through extended test-time thinking.
By Dmitrii Kharlapenko, Terry Jingchen Zhang, Arth Singh, Alessandro Stolfo, Arthur Conmy, Mrinmaya Sachan, Zhijing Jin
A*-Thought-V2 is a framework that models Chain-of-Thought reasoning as a geometric trajectory in a 3D PCA space, using explicit-implicit latent tokens to compress steps that deviate from the main question-to-solution direction. The method measures alignment angles to decide which steps remain text and which become latent, and introduces stepwise embedding forcing and label forcing to train the architecture. Experiments on Qwen models show up to 2.6% accuracy gains, halved response length, and significant reductions in computation and training time.
By Xiaoang Xu, Siyuan Liu, Shuo Wang, Junlan Feng, Fanyu Meng, Zhu Zhang, Jixun Wang, Xiaorong Wang, Zihan Zhou, Xin Li, Chaojun Xiao, Yiming Zhang, Huijia Wu, Liuyu Xiang, Peipei Li, Zhaofeng He
arXiv:2606. 16360v1 Announce Type: cross Abstract: Chain-of-thought (CoT) prompting improves reasoning in large language models (LLMs) by externalizing intermediate computation as discrete text tokens, but this textual interface also introduces redundancy and inference overhead.
By Hanyu Lin, Min Cai, Jiawei Wen, Haodi Zhang
arXiv:2606. 15733v1 Announce Type: cross Abstract: Instruction-tuned language models can answer the same causal-reasoning question differently after its English variable names are replaced by type-preserving placeholders, although the structural causal model and the gold answer are unchanged.
By Zhenyu Yu
arXiv:2606. 05194v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly being deployed to make decisions that require trading off near-term gains against long-term consequences, yet little is known about how they internally represent or resolve these tradeoffs.
By Ian Rios-Sialer, Shantanu Darveshi, Shuai Jiang, Avigya Paudel, Anastasiia Pronina, Ipshita Bandyopadhyay, Justin Shenk
arXiv:2609.38806v1 Announce Type: cross
Abstract: Next-token prediction has driven remarkable progress in large language models, yet a growing body of evidence suggests that they can struggle on prob...
By Woosang Jeon, Jaeyeon Kim, Sham Kakade, Yilun Du, Amrit Singh Bedi, Arun Kumar Chithanar, Chul Lee, Taehyeong Kim, Sitan Chen
The paper investigates how large language models perform propositional logical reasoning by conducting a causal mechanistic analysis on the PropLogic-MI benchmark. It identifies four interlocking mechanisms—Staged Computation, Information Transmission, Fact Retrospection, and Specialized Attention Heads—that organize the reasoning process across layers. The study demonstrates that these mechanisms recur across different model families, rule categories, and reasoning hops, indicating a structured, layer‑organized internal process for propositional reasoning.
By Danchun Chen, Qiyao Yan, Chenpeng Wang, Liangming Pan
arXiv:2607. 18100v1 Announce Type: new Abstract: Extended reasoning has become standard for frontier Large Language Models (LLMs), yet the trajectories these models produce remain largely uncontrollable.
By Sheldon Yu, Tong Yu, Xunyi Jiang, Rohan Surana, Gagan Mundada, Sungchul Kim, Lina Yao, Julian McAuley, Junda Wu