The monograph explores the relationships between Gaussian processes and reproducing kernel Hilbert spaces (RKHS), two widely used approaches that rely on positive definite kernels. It examines how these frameworks connect and are equivalent across key topics such as regression, interpolation, numerical integration, distributional discrepancies, statistical dependence, and Gaussian process sample path properties. By establishing a unifying perspective based on the equivalence between the Gaussian Hilbert space and the RKHS, the work aims to bridge methods developed independently by the machine learning, statistics, and numerical analysis communities.
By Motonobu Kanagawa, Philipp Hennig, Dino Sejdinovic, Bharath K. Sriperumbudur
arXiv:2209. 14125v3 Announce Type: replace-cross Abstract: Diffusion models have proven to be a flexible and effective framework for modelling probability distributions on finite-dimensional spaces.
By Angus Phillips, Thomas Seror, Michael Hutchinson, Valentin De Bortoli, Arnaud Doucet, Emile Mathieu
arXiv:2605. 10285v2 Announce Type: replace-cross Abstract: We present a theoretically grounded Gaussian process framework that leverages neural feature maps to construct expressive kernels.
By Anthony Stephenson
arXiv:2512. 06143v2 Announce Type: replace Abstract: Despite a large corpus of recent work on scaling up Gaussian processes, a stubborn trade-off between computational speed, prediction and uncertainty quantification accuracy, and customizability persists.
By Marcus M. Noack, Mark D. Risser, Hengrui Luo, Vardaan Tekriwal, Ronald J. Pandolfi
arXiv:2607. 02199v1 Announce Type: cross Abstract: Mutual information (MI)-inspired feature learning techniques are capable of generating low-dimensional embeddings that retain nonlinear dependence structures, but direct estimations of MI suffer from noisy probability distribution estimates in the low-data regime.
By Preston Pitzer, Anish Pradhan, Harpreet S. Dhillon
arXiv:2608.29265v1 Announce Type: cross
Abstract: Kernel density estimation (KDE) is one of the most fundamental statistical estimators of density functions. Its direct implementation on a dataset of...
By Xie Wang, Nicolas Langren\'e, Wen Chen
arXiv:2607. 18282v1 Announce Type: new Abstract: Bayesian Optimization is widely used for expensive black-box optimization, yet its success often depends on choosing a kernel that matches the objective's unknown structure.
By Weibo Huang, Cheng Hua
arXiv:2608. 08704v1 Announce Type: cross Abstract: Kernel spectral clustering with a single bandwidth can be inadequate for data exhibiting multiple characteristic pairwise-distance scales, a problem particularly prevalent in the high-dimensional regime.
By Zeqin Lin, Guangming Pan, Zhixiang Zhang, Yinbing Zhou
The paper introduces an amortized learning framework for selecting bandwidths in kernel density estimation by optimizing the logarithmic score across a distribution of tasks. It uses a truncated-and-renormalized bounded-support formulation and affine standardization to achieve stable learning and transferability across different intervals. Experiments on Gaussian samples, a multi-family benchmark, and randomized Gaussian mixtures demonstrate that the learned selector outperforms traditional methods such as Silverman’s rule, Sheather–Jones, and least‑squares cross‑validation, especially for small or heterogeneous samples.
By Junyi Liang, Hailiang Du
The paper introduces the Sparse Landmark Embedding (SLE) kernel, a new framework that removes the need for conditionally negative definite (CND) distance measures in kernel methods and Gaussian Processes. By embedding each input into a sparse feature vector using compactly supported bump functions centered at all training points, any standard positive semi-definite (PSD) kernel can be applied in this embedding space, guaranteeing PSD for arbitrary distance measures. The authors provide theoretical guarantees on PSD, sparsity, stability, and universal approximation, and show through experiments with geodesic and Wasserstein distances that the SLE kernel matches or surpasses domain-specific baselines in predictive accuracy and uncertainty quantification.
By Marcus M. Noack, Maher B. Alghalayini, Mark D. Risser
The paper argues against using the Gaussian (squared exponential/RBF) kernel as a default in Gaussian process regression, citing its brittleness. It shows that the kernel leads to unrealistically small conditional variances, causing overconfidence in predictive uncertainty, and that this small variance induces numerical ill‑conditioning, necessitating tricks like nugget terms that alter the model. The authors attribute these issues to the kernel’s analytic, highly smooth nature and suggest that analytic stationary kernels in general should be avoided.
arXiv:2501. 10870v2 Announce Type: replace-cross Abstract: The principal objective of this work is twofold within nonparametric regression settings: (1) to establish the minimax optimal convergence rates for fixed-bandwidth Gaussian kernel spectral algorithms when the true regression function resides in a Sobolev space, and (2) to apply Gaussian spectral algorithms for achieving robust and adaptive transfer learning under concept shift.
By Haotian Lin, Matthew Reimherr