arXiv Machine Learning By Binbin Lin, Wei Chen, Yalun Li, Wenxiao Wang, Jieping Ye, Xiaofei He

From Self-Attention to Connection Laplacian: A Unified Operator View of Transformers

Read the original on arXiv Machine Learning →

arXiv:2607. 10677v1 Announce Type: new Abstract: Self-attention is a ubiquitous primitive in modern sequence models, yet its operator-level geometry is only partially understood.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
1d ago

Let the Heads Talk: Beyond Diagonal Graph Attention

The paper introduces Topological Attention (Top‑A), a multi‑head attention mechanism that extends standard diagonal edge maps by allowing off‑diagonal, edge‑conditioned communication across attention heads. By isolating the transport primitive through quiver representations, the authors show that standard multi‑head attention only implements diagonal edge maps, whereas Top‑A learns additional cross‑head routes while preserving the original same‑head paths. Experiments on relational reasoning, heterogeneous graph learning, and algorithmic reasoning demonstrate that cross‑head transport is most beneficial when tasks require interaction‑dependent transformations, whereas heterophily alone does not provide a systematic advantage.

By Riccardo Ali, Alessio Borgi, Mario Severino, Alessio Gravina, Davide Bacciu, Pietro Li\`o, Christopher Irwin
arXiv Machine Learning
Aug 20

The Diffusion-Attention Connection

arXiv:2604. 09560v2 Announce Type: replace Abstract: Softmax attention is the row-normalized operator of a diffusion map: both normalize a learned score into a Markov operator, and differ only in what the score is allowed to contain.

By Julio Candanedo