arXiv Machine Learning By Kailen Hargenrader, Edoardo Calvello, Bohan Chen

Attention Kernels for Learning Maps Between Heavy-Tailed Measures

Read the original on arXiv Machine Learning →

The paper introduces attention kernels that replace the exponential function in transformer softmax to better handle operator learning on probability measures with heavy-tailed (polynomial) distributions. Two new benchmarks with closed‑form targets are constructed to evaluate how different kernel growth rates and data preprocessing affect performance. The study finds that slower‑growing kernels prevent ensemble collapse on heavy‑tailed tasks, while softmax with symlog preprocessing only succeeds on a subset of problems, and that all kernels perform similarly on Gaussian data.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 4

Entropy-Generated Attention Beyond Softmax and Entmax: Kaniadakis and Reciprocal-Symmetric Abe Operators

The paper introduces two novel attention operators derived from generalized statistical entropies. The Kaniadakis entropy yields a full-support normalization with algebraically decaying weights, while the Abe entropy produces an implicit reciprocal-symmetric operator. The authors analyze these operators through a Fisher-metric Lagrangian framework, compare them to Softmax and entmax, and provide a tangent-gradient test to distinguish changes in attention profiles from mere scaling effects.

By Gunn Kim