arXiv Machine Learning

Entropy-Generated Attention Beyond Softmax and Entmax: Kaniadakis and Reciprocal-Symmetric Abe Operators

The paper introduces two novel attention operators derived from generalized statistical entropies. The Kaniadakis entropy yields a full-support normalization with algebraically decaying weights, while the Abe entropy produces an implicit reciprocal-symmetric operator. The authors analyze these operators through a Fisher-metric Lagrangian framework, compare them to Softmax and entmax, and provide a tangent-gradient test to distinguish changes in attention profiles from mere scaling effects.

arXiv Machine Learning
1d ago

Attention Kernels for Learning Maps Between Heavy-Tailed Measures

The paper introduces attention kernels that replace the exponential function in transformer softmax to better handle operator learning on probability measures with heavy-tailed (polynomial) distributions. Two new benchmarks with closed‑form targets are constructed to evaluate how different kernel growth rates and data preprocessing affect performance. The study finds that slower‑growing kernels prevent ensemble collapse on heavy‑tailed tasks, while softmax with symlog preprocessing only succeeds on a subset of problems, and that all kernels perform similarly on Gaussian data.

By Kailen Hargenrader, Edoardo Calvello, Bohan Chen
arXiv AI
Jul 23

Geometric Attention: A Regime-Explicit Operator Semantics for Transformer Attention

arXiv:2601. 11618v2 Announce Type: replace-cross Abstract: Geometric Attention (GA) specifies an attention layer by four independent inputs: a finite carrier (what indices are addressable), an evidence-kernel rule (how masked proto-scores and a link induce nonnegative weights), a probe family (which observables are treated as admissible), and an anchor/update rule (which representative kernel is selected and how it is applied).

By Luis Rosario Freytes
arXiv Machine Learning
Aug 20

The Diffusion-Attention Connection

arXiv:2604. 09560v2 Announce Type: replace Abstract: Softmax attention is the row-normalized operator of a diffusion map: both normalize a learned score into a Markov operator, and differ only in what the score is allowed to contain.

By Julio Candanedo