arXiv Machine Learning

From Self-Attention to Connection Laplacian: A Unified Operator View of Transformers

arXiv:2607. 10677v1 Announce Type: new Abstract: Self-attention is a ubiquitous primitive in modern sequence models, yet its operator-level geometry is only partially understood.

arXiv Machine Learning
1d ago

Let the Heads Talk: Beyond Diagonal Graph Attention

The paper introduces Topological Attention (Top‑A), a multi‑head attention mechanism that extends standard diagonal edge maps by allowing off‑diagonal, edge‑conditioned communication across attention heads. By isolating the transport primitive through quiver representations, the authors show that standard multi‑head attention only implements diagonal edge maps, whereas Top‑A learns additional cross‑head routes while preserving the original same‑head paths. Experiments on relational reasoning, heterogeneous graph learning, and algorithmic reasoning demonstrate that cross‑head transport is most beneficial when tasks require interaction‑dependent transformations, whereas heterophily alone does not provide a systematic advantage.

By Riccardo Ali, Alessio Borgi, Mario Severino, Alessio Gravina, Davide Bacciu, Pietro Li\`o, Christopher Irwin
arXiv Machine Learning
Aug 20

The Diffusion-Attention Connection

arXiv:2604. 09560v2 Announce Type: replace Abstract: Softmax attention is the row-normalized operator of a diffusion map: both normalize a learned score into a Markov operator, and differ only in what the score is allowed to contain.

By Julio Candanedo
arXiv Machine Learning
4d ago

Lost in Tokenization: Fundamental Trade-offs in Graph Tokenization for Transformers

The paper investigates how the choice of graph tokenization affects transformer expressivity. It analyzes three tokenization families—spectral, random‑walk, and adjacency—showing that each induces different depth requirements and that some tokenizations are inherently lossy or ill‑conditioned for certain tasks. The authors prove lower bounds and impossibility results for converting between tokenizations and validate these findings with experiments on synthetic and real‑world data.

By Maya Bechler-Speicher, Gilad Yehudai, Gil Harari, Clayton Sanford, Amir Globerson, Joan Bruna
Hugging Face Trending Papers
Jul 13

Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks

We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous works on Transformer learning dynamics have so far been mostly tied to specific tasks, we study a generalized class of inductive tasks that unifies several synthetic tasks known in the literature, including in-context n-grams and multi-hop reasoning.