arXiv Machine Learning

Stability of Measure-to-Measure Transformers on Sub-Gaussian Data

This paper provides a mathematical analysis of measure-to-measure transformers, showing that they map sub‑Gaussian inputs to sub‑Gaussian outputs and are Hölder continuous with respect to the 1‑Wasserstein distance on suitable spaces. It establishes error‑propagation estimates for transformers applied to empirical approximations of sub‑Gaussian data and investigates a mean‑field analogue of cross‑attention, revealing distinct Hölder regularity and sample‑complexity for its two inputs. The results culminate in approximation guarantees for measure‑to‑measure transformers, offering a rigorous stability and finite‑sample theory for transformers on sub‑Gaussian data.

Hugging Face Trending Papers
Jul 30

Generalization Bounds on Optimal Control for Transformer Training and Wasserstein Distributional Robustness

We derive finite-sample generalization bounds for Transformers trained with dynamic programming recursions. Building on the doubly lifted, measure-valued formulation of Transformer dynamics, we view data sets as probability laws on pairs of empirical input-output measures, allowing us to interpret the training problem as a finite-horizon Markovian control problem.

arXiv Machine Learning
5d ago

Attention Kernels for Learning Maps Between Heavy-Tailed Measures

The paper introduces attention kernels that replace the exponential function in transformer softmax to better handle operator learning on probability measures with heavy-tailed (polynomial) distributions. Two new benchmarks with closed‑form targets are constructed to evaluate how different kernel growth rates and data preprocessing affect performance. The study finds that slower‑growing kernels prevent ensemble collapse on heavy‑tailed tasks, while softmax with symlog preprocessing only succeeds on a subset of problems, and that all kernels perform similarly on Gaussian data.

By Kailen Hargenrader, Edoardo Calvello, Bohan Chen
arXiv Machine Learning
Sep 2

Performance-Efficiency Tradeoffs in Transformers: An Approximation Theory Perspective

The paper examines how to allocate attention heads and head dimensions across Transformer layers to balance expressivity and efficiency. It provides a mathematical analysis of early layers’ role in information extraction and characterizes the trade‑off between head count and dimension under a fixed parameter budget. The authors prove a saturation effect of softmax activations, showing that increasing head dimensions yields diminishing returns, especially for long sequences, and propose strategies for efficient parameter allocation across layers.

By Ruoxi Yu, Haotian Jiang, Jingpu Cheng, Penghao Yu, Qianxiao Li, Zhong Li