arXiv:2607. 00479v1 Announce Type: new Abstract: Transformer-based large models have demonstrated remarkable generalization abilities across different tasks by leveraging a context-aware attention module for in-context learning.
By Peilin Liu, Ding-Xuan Zhou
arXiv:2606. 27748v1 Announce Type: cross Abstract: Transformer models rely on attention mechanism to capture long-range dependencies but suffer from quadratic complexity, limiting their scalability to long sequences.
By Haoran Zhang, Feng Zhou
arXiv:2610.03201v1 Announce Type: cross
Abstract: Gaussian Processes (GPs) are a powerful tool for modelling and quantifying uncertainty in functional relationships. However, they require practitione...
By Callum Lau, Jeremias Knoblauch, Louis Sharrock
arXiv:2605. 10285v2 Announce Type: replace-cross Abstract: We present a theoretically grounded Gaussian process framework that leverages neural feature maps to construct expressive kernels.
By Anthony Stephenson
This paper provides a mathematical analysis of measure-to-measure transformers, showing that they map sub‑Gaussian inputs to sub‑Gaussian outputs and are Hölder continuous with respect to the 1‑Wasserstein distance on suitable spaces. It establishes error‑propagation estimates for transformers applied to empirical approximations of sub‑Gaussian data and investigates a mean‑field analogue of cross‑attention, revealing distinct Hölder regularity and sample‑complexity for its two inputs. The results culminate in approximation guarantees for measure‑to‑measure transformers, offering a rigorous stability and finite‑sample theory for transformers on sub‑Gaussian data.
By Frank Cole, Nicholas H. Nelsen, Takashi Furuya
arXiv:2607. 23869v1 Announce Type: cross Abstract: Randomized features provide a scalable approximation to kernel machines, but their performance depends strongly on the choice of feature distribution.
By Masoud Badiei Khuzani, Sharath Honnaiah, Atiq Islam, Alex Cozzi, Abraham Bagherjeiran
arXiv:2605. 08475v3 Announce Type: replace-cross Abstract: In this paper, we study in-context kernel ridge regression (KRR) with Gaussian kernels and show, both theoretically and empirically, that a standard softmax-attention transformer can approximate the KRR predictor during its forward pass.
By Mingsong Yan, Dongyang Li, Charles Kulick, Sui Tang
arXiv:2605. 18848v3 Announce Type: replace Abstract: This paper introduces Exact Linear Attention (ELA), a mechanism that achieves linear computational complexity for Transformer attention by exploiting the exact decomposition property of kernel functions, thereby eliminating approximation error.
By Weinuo Ou
arXiv:2608. 09558v1 Announce Type: new Abstract: How expressive is prompting a transformer?
By Alexander Hsu, Rongjie Lai
arXiv:2602. 23006v2 Announce Type: replace-cross Abstract: Simulating a Gaussian process requires sampling from a high-dimensional Gaussian distribution, which scales cubically with the number of sample locations.
By Arsalan Jawaid, Abdullah Karatas, J\"org Seewig
arXiv:2605. 11287v2 Announce Type: replace-cross Abstract: A persistent paradox in time-series forecasting is that structurally simple MLP and linear models often outperform high-capacity Transformers.
By Jevon Twitty, Vinh Pham, Nitiwith Rotchanarak, Viresh Pati, Yubin Kim, Shihao Yang, Jiecheng Lu
arXiv:2606. 13694v1 Announce Type: cross Abstract: Mobile sleep staging serves as a foundational infrastructure for in-home sleep monitoring and closed-loop modulation.
By Guisong Liu, Pengfei Wei, Jainsong Zhang, Martin Dresler