The paper introduces attention kernels that replace the exponential function in transformer softmax to better handle operator learning on probability measures with heavy-tailed (polynomial) distributions. Two new benchmarks with closed‑form targets are constructed to evaluate how different kernel growth rates and data preprocessing affect performance. The study finds that slower‑growing kernels prevent ensemble collapse on heavy‑tailed tasks, while softmax with symlog preprocessing only succeeds on a subset of problems, and that all kernels perform similarly on Gaussian data.
By Kailen Hargenrader, Edoardo Calvello, Bohan Chen
arXiv:2608. 11173v1 Announce Type: cross Abstract: The attention mechanism forms the foundation of many modern AI models such as the Transformer.
By Eric A. F. Reinhardt, Adam J. Hauser
arXiv:2601. 11618v2 Announce Type: replace-cross Abstract: Geometric Attention (GA) specifies an attention layer by four independent inputs: a finite carrier (what indices are addressable), an evidence-kernel rule (how masked proto-scores and a link induce nonnegative weights), a probe family (which observables are treated as admissible), and an anchor/update rule (which representative kernel is selected and how it is applied).
By Luis Rosario Freytes
arXiv:2508. 17821v3 Announce Type: replace-cross Abstract: This paper investigates the limitations of the normalization in attention mechanisms.
By Timur Mudarisov, Mikhail Burtsev, Tatiana Petrova, Radu State
arXiv:2604. 09560v2 Announce Type: replace Abstract: Softmax attention is the row-normalized operator of a diffusion map: both normalize a learned score into a Markov operator, and differ only in what the score is allowed to contain.
By Julio Candanedo
arXiv:2606. 12059v1 Announce Type: new Abstract: We address transformer attention on energy-constrained physical substrates.
By Fabio Pasqualetti, Taosha Guo