arXiv AI By Tiberiu Musat, Tiago Pimentel, Nicholas Zucchet, Thomas Hofmann

Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks

Read the original on arXiv AI →

arXiv:2607. 11875v1 Announce Type: cross Abstract: We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 13

Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks

We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous works on Transformer learning dynamics have so far been mostly tied to specific tasks, we study a generalized class of inductive tasks that unifies several synthetic tasks known in the literature, including in-context n-grams and multi-hop reasoning.

arXiv Machine Learning
Sep 17

Understanding the Staged Dynamics of Transformers in Learning Latent Structure

The paper investigates how transformer models learn latent structure by training a small decoder-only transformer on three variants of the Alchemy benchmark. It finds that the model acquires different components of latent structure in discrete stages, with a notable asymmetry: it robustly composes fundamental transitions but struggles to decompose complex examples into atomic transitions. Layer‑specific causal interventions reveal plasticity windows where freezing layers delays or prevents stage completion, offering a detailed view of capability evolution during training.

By Rohan Saha, Farzane Aminmansour, Alona Fyshe
arXiv Computation and Language
Aug 28

Why Knowing Both Hops Is Not Enough: Understanding Two-Hop Generalization in Language Models

Large language models can solve complex multi‑hop tasks but often fail on simple two‑hop queries, even when each hop is individually correct. In a controlled symbolic setting, the authors find that models generalize reliably when the second hop follows the training distribution, but always fail when it deviates. Mechanistic analysis shows that successful generalization relies on consistent intermediate representations across contexts, whereas failures arise from a mismatch between lower‑layer representation construction and upper‑layer mapping to outputs. The study proposes a recurrent‑style training strategy that improves out‑of‑distribution two‑hop generalization.

By Zili Zhang, Yilin Wang, Heng Wang, Herun Wan, Minnan Luo