arXiv:2606. 12318v1 Announce Type: cross Abstract: Neural operators approximate mappings between function spaces, but often generalize poorly to other operators and usually require fine-tuning or retraining.
By Minghui Yang, Ling Guo, Liu Yang
The paper proposes that large language models (LLMs) encode high‑level concepts as linear directions within their activation space and that they can use subspaces and vector algebra to perform tasks. By analyzing functional modules and residual streams during in‑context learning (ICL), the authors find that LLMs can create evidence‑accumulating subspaces and solve ICL tasks through simple algebraic operations within those subspaces.
By Jung H. Lee, Sujith Vijayan
arXiv:2607. 26775v1 Announce Type: new Abstract: Many kinds of data have structure along one or more axes: words in a sentence, pixels in an image, nodes in a tree, frames in audio, or cells in a 3D volume.
By Mahesh Godavarti
arXiv:2602. 24264v2 Announce Type: replace-cross Abstract: Compositional generalization, the ability to recognize familiar parts in novel contexts, is a defining property of intelligent systems.
By Arnas Uselis, Andrea Dittadi, Seong Joon Oh
arXiv:2602.02611v2 Announce Type: replace
Abstract: A prevailing paradigm in modern representation learning is the map-first approach, in which a representation map is learned from reconstruction, em...
By David Vigouroux (ANITI, IMT Atlantique - DSD, LaTIM), Lucas Drumetz (IMT Atlantique - MEE, Lab-STICC\_OSE, ODYSSEY), Ronan Fablet (IMT Atlantique - MEE, Lab-STICC\_OSE, ODYSSEY), Fran\c{c}ois Rousseau (IMT Atlantique - DSD, LaTIM)
arXiv:2505.09716v3 Announce Type: replace-cross
Abstract: Out-of-distribution (OOD) generalisation is considered a hallmark of human and animal intelligence. To achieve OOD through composition, a sys...
By George Dimitriadis, Spyridon Samothrakis
The paper investigates how transformer language models perform few‑shot learning for a simple addition task, showing that the ability is concentrated in a handful of attention heads. Using dimensionality reduction, the authors identify low‑dimensional subspaces—three heads with six‑dimensional spaces in Llama‑3‑8B‑Instruct—where specific dimensions encode the units digit via trigonometric patterns and magnitude via low‑frequency components. They also derive a mathematical identity linking aggregator and extractor subspaces, enabling tracking of information flow from examples to the final prediction.
By Xinyan Hu, Kayo Yin, Michael I. Jordan, Jacob Steinhardt, Lijie Chen
arXiv:2512. 12225v3 Announce Type: replace Abstract: Developing artificial agents that unify representation, memory, adaptation, and prediction remains a fundamental challenge in artificial intelligence.
By Laha Ale
arXiv:2602. 03282v2 Announce Type: replace-cross Abstract: A common assumption in representation learning is that globally well-distributed embeddings support robust and generalizable representations.
By Jiwan Chung, Seon Joo Kim
We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous works on Transformer learning dynamics have so far been mostly tied to specific tasks, we study a generalized class of inductive tasks that unifies several synthetic tasks known in the literature, including in-context n-grams and multi-hop reasoning.
arXiv:2507. 11688v4 Announce Type: replace Abstract: Contemporary large models often exhibit behaviors suggesting the presence of low-level primitives that compose into modules with richer functionality, but these fundamental building blocks remain poorly understood.
By Travis Pence, Daisuke Yamada, Vikas Singh
arXiv:2601. 22402v2 Announce Type: replace-cross Abstract: Rotary Positional Embeddings (RoPE) have become the standard for Large Language Models (LLMs) due to their ability to encode relative positions through geometric rotation.
By Kanishk Awadhiya