The paper investigates how transformer models learn latent structure by training a small decoder-only transformer on three variants of the Alchemy benchmark. It finds that the model acquires different components of latent structure in discrete stages, with a notable asymmetry: it robustly composes fundamental transitions but struggles to decompose complex examples into atomic transitions. Layer‑specific causal interventions reveal plasticity windows where freezing layers delays or prevents stage completion, offering a detailed view of capability evolution during training.
By Rohan Saha, Farzane Aminmansour, Alona Fyshe
arXiv:2604. 22951v2 Announce Type: replace Abstract: Natural language data follows a power-law distribution, with most knowledge and skills appearing at very low frequency.
By Zixuan Wang, Xingyu Dang, Jason D. Lee, Kaifeng Lyu
arXiv:2607. 17166v1 Announce Type: new Abstract: Transformer-based large language models (LLMs) continue to achieve state-of-the-art performance across various natural language processing tasks.
By Luyu Qiu, Jianing Li, Hwanhee Kim, Xiaoyong Wei, Yueyuan Zheng, Janet Hsiao, Lei Chen
arXiv:2602. 22600v2 Announce Type: replace-cross Abstract: Training selects for behavior, not circuitry: many weight configurations can implement the same function.
By Joshua S. Schiffman
arXiv:2607. 25663v1 Announce Type: new Abstract: Transformer adaptation is typically distributed across model depth, even when the intended change is narrow.
By Rebecca Ramnauth, Brian Scassellati
We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous works on Transformer learning dynamics have so far been mostly tied to specific tasks, we study a generalized class of inductive tasks that unifies several synthetic tasks known in the literature, including in-context n-grams and multi-hop reasoning.
arXiv:2602. 14872v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has been a main driver of recent breakthroughs in large reasoning models.
By Yu Huang, Zixin Wen, Yuejie Chi, Yuting Wei, Aarti Singh, Yingbin Liang, Yuxin Chen
arXiv:2607. 11875v1 Announce Type: cross Abstract: We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models.
By Tiberiu Musat, Tiago Pimentel, Nicholas Zucchet, Thomas Hofmann
The paper investigates why large language models (LLMs) fail to maintain continuous mixtures of token embeddings—used in latent-state reasoning—to preserve multiple reasoning paths. Through theory and experiments, it identifies three failure sources: transformer geometry distortion, amplification or contraction dynamics from softmax and autoregressive feedback, and the need for context-dependent corrections that scale with mixture size. Empirical results confirm the predicted transition between contraction and amplification and show pretrained models largely fall on the amplifying side.
By Ali Backour
arXiv:2605. 04970v3 Announce Type: replace-cross Abstract: Modern LLMs show mastery over an ever-growing range of skills, as well as the ability to compose them flexibly.
By Antonin Berthon, Nicolas Astorga, Mihaela van der Schaar
The paper investigates how large language models learn new tasks in-context, comparing rule-based instruction following to example-based few-shot prompting across five diverse tasks. Results show that models generally learn more reliably from rule descriptions than from examples alone, and adding more examples does not consistently improve performance. Instruction tuning further enhances rule-based learning while preserving example-based capabilities, with rule advantages being strongest for algebraic tasks and weaker for tasks requiring distributional sensitivity or parametric knowledge.
By Xiang Fu, Seungmin Cho, Yukyung Lee, Najoung Kim
arXiv:2602. 02470v2 Announce Type: replace Abstract: Autoregressive large language models (LLMs) have achieved remarkable success in many complex tasks, yet they can still fail in very simple logical reasoning such as the "reversal curse" -- when trained on forward knowledge data of the form "$A \rightarrow B$" (e.
By Xutao Ma, Yixiao Huang, Hanlin Zhu, Somayeh Sojoudi