Random Feature Gaussian Process Attention: Linear-Time Probabilistic Attention with Calibrated Uncertainty
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2607. 00479v1 Announce Type: new Abstract: Transformer-based large models have demonstrated remarkable generalization abilities across different tasks by leveraging a context-aware attention module for in-context learning.
arXiv:2606. 27748v1 Announce Type: cross Abstract: Transformer models rely on attention mechanism to capture long-range dependencies but suffer from quadratic complexity, limiting their scalability to long sequences.
arXiv:2610.03201v1 Announce Type: cross Abstract: Gaussian Processes (GPs) are a powerful tool for modelling and quantifying uncertainty in functional relationships. However, they require practitione...
arXiv:2605. 10285v2 Announce Type: replace-cross Abstract: We present a theoretically grounded Gaussian process framework that leverages neural feature maps to construct expressive kernels.
This paper provides a mathematical analysis of measure-to-measure transformers, showing that they map sub‑Gaussian inputs to sub‑Gaussian outputs and are Hölder continuous with respect to the 1‑Wasserstein distance on suitable spaces. It establishes error‑propagation estimates for transformers applied to empirical approximations of sub‑Gaussian data and investigates a mean‑field analogue of cross‑attention, revealing distinct Hölder regularity and sample‑complexity for its two inputs. The results culminate in approximation guarantees for measure‑to‑measure transformers, offering a rigorous stability and finite‑sample theory for transformers on sub‑Gaussian data.
arXiv:2607. 23869v1 Announce Type: cross Abstract: Randomized features provide a scalable approximation to kernel machines, but their performance depends strongly on the choice of feature distribution.