The paper introduces Topological Attention (Top‑A), a multi‑head attention mechanism that extends standard diagonal edge maps by allowing off‑diagonal, edge‑conditioned communication across attention heads. By isolating the transport primitive through quiver representations, the authors show that standard multi‑head attention only implements diagonal edge maps, whereas Top‑A learns additional cross‑head routes while preserving the original same‑head paths. Experiments on relational reasoning, heterogeneous graph learning, and algorithmic reasoning demonstrate that cross‑head transport is most beneficial when tasks require interaction‑dependent transformations, whereas heterophily alone does not provide a systematic advantage.
By Riccardo Ali, Alessio Borgi, Mario Severino, Alessio Gravina, Davide Bacciu, Pietro Li\`o, Christopher Irwin
arXiv:2604. 09560v2 Announce Type: replace Abstract: Softmax attention is the row-normalized operator of a diffusion map: both normalize a learned score into a Markov operator, and differ only in what the score is allowed to contain.
By Julio Candanedo
Positional encodings (PEs) are essential for Transformers. Yet designing effective PEs for non-Euclidean graphs remains challenging.
arXiv:2608. 01283v1 Announce Type: new Abstract: All Transformer-based large language models compute attention via the Euclidean inner product, an architectural choice that Dong et al.
By Sen Song
arXiv:2602.15239v3 Announce Type: replace
Abstract: Transformers have achieved remarkable success across domains, motivating the rise of Graph Transformers (GTs) as attention-based architectures for...
By Javier Porras-Valenzuela, Zhiyang Wang, Teresa Shang, Yusu Wang, Alejandro Ribeiro
arXiv:2606. 25293v1 Announce Type: new Abstract: Positional encodings (PEs) are essential for Transformers.
By Yipeng Zhang, Zhongtian Sun, Pietro Li\`o, Kelin Xia