The paper investigates whether large language models (LLMs) maintain belief states—probability distributions over latent variables—by embedding a controllable latent variable into natural text. An LLM teacher generates ordinary text while subtly steering it along one of eight sparse autoencoder directions, which follow a ring-shaped Markov chain. A small transformer trained on this data successfully tracks the Bayesian posterior of the planted variable and arranges the eight states on a ring in the same order as the Markov chain, suggesting a link between concept geometry and latent variable dynamics.
arXiv:2511. 21594v3 Announce Type: replace Abstract: Large language models (LLMs) achieve state-of-the-art results across many natural language tasks, but their internal mechanisms remain difficult to interpret.
By Alex Ning, Vainateya Rangaraju, Yen-Ling Kuo
arXiv:2511.19328v3 Announce Type: replace
Abstract: Language modeling has shown us that transformers can discover latent structure from context, but the dynamics of how they acquire different compone...
By Rohan Saha, Farzane Aminmansour, Alona Fyshe
Large language models (LLMs) trained on next-token prediction exhibit remarkable in-context learning (ICL) abilities, yet the representations that support ICL remain poorly understood. We consider suc...
arXiv:2609.17376v1 Announce Type: new
Abstract: Large language models (LLMs) trained on next-token prediction exhibit remarkable in-context learning (ICL) abilities, yet the representations that supp...
By Daniel Balcells, Andrew Jun Lee, Chirag Rastogi, Paul M. Riechers, Adam Shai, Xavier Poncini
A*-Thought-V2 is a framework that models Chain-of-Thought reasoning as a geometric trajectory in a 3D PCA space, using explicit-implicit latent tokens to compress steps that deviate from the main question-to-solution direction. The method measures alignment angles to decide which steps remain text and which become latent, and introduces stepwise embedding forcing and label forcing to train the architecture. Experiments on Qwen models show up to 2.6% accuracy gains, halved response length, and significant reductions in computation and training time.
By Xiaoang Xu, Siyuan Liu, Shuo Wang, Junlan Feng, Fanyu Meng, Zhu Zhang, Jixun Wang, Xiaorong Wang, Zihan Zhou, Xin Li, Chaojun Xiao, Yiming Zhang, Huijia Wu, Liuyu Xiang, Peipei Li, Zhaofeng He