arXiv Machine Learning

Using Composition Operators to Linearize LLM Semantic Transformations

arXiv AI
3d ago

Functional Subspace, where language models can use vector algebra to solve problems

The paper proposes that large language models (LLMs) encode high‑level concepts as linear directions within their activation space and that they can use subspaces and vector algebra to perform tasks. By analyzing functional modules and residual streams during in‑context learning (ICL), the authors find that LLMs can create evidence‑accumulating subspaces and solve ICL tasks through simple algebraic operations within those subspaces.

By Jung H. Lee, Sujith Vijayan
arXiv Machine Learning
Sep 23

Discovering Data Manifold Geometry through Geometric Properties

arXiv:2602.02611v2 Announce Type: replace Abstract: A prevailing paradigm in modern representation learning is the map-first approach, in which a representation map is learned from reconstruction, em...

By David Vigouroux (ANITI, IMT Atlantique - DSD, LaTIM), Lucas Drumetz (IMT Atlantique - MEE, Lab-STICC\_OSE, ODYSSEY), Ronan Fablet (IMT Atlantique - MEE, Lab-STICC\_OSE, ODYSSEY), Fran\c{c}ois Rousseau (IMT Atlantique - DSD, LaTIM)
arXiv AI
Sep 21

Understanding In-context Learning of Addition via Activation Subspaces

The paper investigates how transformer language models perform few‑shot learning for a simple addition task, showing that the ability is concentrated in a handful of attention heads. Using dimensionality reduction, the authors identify low‑dimensional subspaces—three heads with six‑dimensional spaces in Llama‑3‑8B‑Instruct—where specific dimensions encode the units digit via trigonometric patterns and magnitude via low‑frequency components. They also derive a mathematical identity linking aggregator and extractor subspaces, enabling tracking of information flow from examples to the final prediction.

By Xinyan Hu, Kayo Yin, Michael I. Jordan, Jacob Steinhardt, Lijie Chen
Hugging Face Trending Papers
Jul 13

Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks

We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous works on Transformer learning dynamics have so far been mostly tied to specific tasks, we study a generalized class of inductive tasks that unifies several synthetic tasks known in the literature, including in-context n-grams and multi-hop reasoning.

arXiv Machine Learning
Jun 11

Composing Linear Layers from Irreducibles

arXiv:2507. 11688v4 Announce Type: replace Abstract: Contemporary large models often exhibit behaviors suggesting the presence of low-level primitives that compose into modules with richer functionality, but these fundamental building blocks remain poorly understood.

By Travis Pence, Daisuke Yamada, Vikas Singh