arXiv Machine Learning By Frank Cole, Nicholas H. Nelsen, Takashi Furuya

Stability of Measure-to-Measure Transformers on Sub-Gaussian Data

Read the original on arXiv Machine Learning →

This paper provides a mathematical analysis of measure-to-measure transformers, showing that they map sub‑Gaussian inputs to sub‑Gaussian outputs and are Hölder continuous with respect to the 1‑Wasserstein distance on suitable spaces. It establishes error‑propagation estimates for transformers applied to empirical approximations of sub‑Gaussian data and investigates a mean‑field analogue of cross‑attention, revealing distinct Hölder regularity and sample‑complexity for its two inputs. The results culminate in approximation guarantees for measure‑to‑measure transformers, offering a rigorous stability and finite‑sample theory for transformers on sub‑Gaussian data.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jul 30

Generalization Bounds on Optimal Control for Transformer Training and Wasserstein Distributional Robustness

We derive finite-sample generalization bounds for Transformers trained with dynamic programming recursions. Building on the doubly lifted, measure-valued formulation of Transformer dynamics, we view data sets as probability laws on pairs of empirical input-output measures, allowing us to interpret the training problem as a finite-horizon Markovian control problem.

arXiv Machine Learning
5d ago

Attention Kernels for Learning Maps Between Heavy-Tailed Measures

The paper introduces attention kernels that replace the exponential function in transformer softmax to better handle operator learning on probability measures with heavy-tailed (polynomial) distributions. Two new benchmarks with closed‑form targets are constructed to evaluate how different kernel growth rates and data preprocessing affect performance. The study finds that slower‑growing kernels prevent ensemble collapse on heavy‑tailed tasks, while softmax with symlog preprocessing only succeeds on a subset of problems, and that all kernels perform similarly on Gaussian data.

By Kailen Hargenrader, Edoardo Calvello, Bohan Chen