The Transformer Revolution, Part 1: Dynamic Processing through Output-Weight Interconnections
arXiv:2608. 03921v2 Announce Type: replace Abstract: This paper offers a new interpretation of the Transformer during inference.
arXiv:2607. 08196v1 Announce Type: new Abstract: As part of a series on first-principles modeling of cognitive functions, this paper attempts to provide a mathematical formulation of thinking and perception.
arXiv:2608. 03921v2 Announce Type: replace Abstract: This paper offers a new interpretation of the Transformer during inference.
arXiv:2604. 02029v2 Announce Type: replace Abstract: Latent space is rapidly emerging as a native substrate for language-based models.
arXiv:2608. 03921v1 Announce Type: new Abstract: This paper offers a new interpretation of the Transformer during inference.
arXiv:2604. 06374v2 Announce Type: replace-cross Abstract: Latent reasoning via continuous chain-of-thoughts (Latent CoT) has emerged as a promising alternative to discrete CoT reasoning.
arXiv:2608. 12398v1 Announce Type: cross Abstract: We propose IM-LEPP (Integrated Multimodal Latent Energy-based Predictive Processing), a hierarchical, energy-based model of multimodal cognition that extends a previously proposed single-modality model (LEPP) to integrate vision and language.
The paper presents a mean‑field analysis of attention in language models, defining an average attention kernel that propagates representations layer by layer. When conditioned on a whole corpus, the kernel predicts the average evolution of representation geometry; when conditioned on a single context, it predicts the expected geometry for that context. The difference between actual attention and the mean‑field prediction—called the mean‑field deviation—captures context‑specific computation, revealing how models diverge from average behavior during training and in few‑shot tasks.
arXiv:2603. 16689v2 Announce Type: replace Abstract: Next-token predictors often appear to develop internal representations of the latent world and its rules.
arXiv:2607. 28942v1 Announce Type: new Abstract: Recently Large Language Models (LLMs) have been increasingly deployed as autonomous agents in applications such as self-reflection, retrieval-augmented generation, and scientific discovery.
The paper presents a mathematical analysis of the Jacobian lens (J‑lens), a method for extracting verbalizable representations from language models. It treats the J‑lens as a first‑order causal transfer operator, showing that its Jacobian matrix serves as an optimal local linear approximation of downstream mappings and that its energy distribution is highly sparse, concentrating in diagonal pathways and critical positions. This sparse, short‑horizon structure explains why the J‑lens can effectively visualize concepts during a model’s reasoning process, and the authors propose a decoupling strategy that further improves its ability to read out correct intermediate concepts.
Latent Recurrent Thoughts (LRT) proposes a method for reasoning with frozen large language models by operating in the model’s continuous representation space. A small auxiliary network generates initial latent vectors, which a tiny recurrent reasoner refines over multiple steps, decoupling computational depth from model size. Experiments on symbolic and natural‑language reasoning tasks show that LRT outperforms prior frozen‑decoder continuous‑space methods and chain‑of‑thought prompting while using far less inference compute.
arXiv:2609.07406v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning improves the reasoning ability of large language models by introducing intermediate computation, but explicit rational...
arXiv:2606. 29971v1 Announce Type: new Abstract: A growing body of work suggests that the reasoning capabilities of large language models are largely latent in their base form, with post-training primarily amplifying rather than introducing them.