arXiv Machine Learning

gp2Scale: A Class of Compactly Supported Non-Stationary Kernels and Distributed Computing for Exact Gaussian Processes on 10 Million Data Points

arXiv:2512. 06143v2 Announce Type: replace Abstract: Despite a large corpus of recent work on scaling up Gaussian processes, a stubborn trade-off between computational speed, prediction and uncertainty quantification accuracy, and customizability persists.

arXiv AI
6d ago

Adaptive multi-resolution Gaussian processes: Scalable exact inference with naturally data-sparse covariance matrices

The paper introduces an adaptive multi‑resolution Gaussian process framework that achieves scalable, exact inference by constructing a naturally data‑sparse covariance matrix using basis functions anchored directly to samples. By shrinking the support domains of these basis functions, the resulting matrix has limited block sizes, ensuring sparsity and enabling efficient computation of its inverse via a sparse Cholesky algorithm. The authors demonstrate that this approach yields exact inference with training cost ≠≠ O(n log^2 n) and prediction cost ≠≠ O(log^d n), while also improving predictive uncertainties through an augmented basis function.

By Yanchuang Cao, Jun Liu, Tengchao Yu, Heng Yong
arXiv Machine Learning
Aug 24

Exact and general decoupled solutions of the LMC Multitask Gaussian Process model

The paper presents an exact, efficient solution for the Linear Model of Co‑regionalization (LMC) multitask Gaussian Process by decoupling latent processes under a mild noise‑model assumption. It introduces a full parametrization of the resulting projected LMC, enabling linear‑time optimization and simplifying tasks such as training updates and leave‑one‑out cross‑validation. Experiments on synthetic and real data demonstrate that projected LMC is competitive with state‑of‑the‑art multitask GP models while offering greater interpretability and computational ease.

By Olivier Truffinet (CEA Saclay), Karim Ammar (CEA Saclay), Jean-Philippe Argaud (EDF R&D), Bertrand Bouriquet (EDF)
arXiv Machine Learning
Sep 17

A General Kernel Framework for Non-CND Distance Measures Using |D|-Dimensional Sparse Landmark Embeddings

The paper introduces the Sparse Landmark Embedding (SLE) kernel, a new framework that removes the need for conditionally negative definite (CND) distance measures in kernel methods and Gaussian Processes. By embedding each input into a sparse feature vector using compactly supported bump functions centered at all training points, any standard positive semi-definite (PSD) kernel can be applied in this embedding space, guaranteeing PSD for arbitrary distance measures. The authors provide theoretical guarantees on PSD, sparsity, stability, and universal approximation, and show through experiments with geodesic and Wasserstein distances that the SLE kernel matches or surpasses domain-specific baselines in predictive accuracy and uncertainty quantification.

By Marcus M. Noack, Maher B. Alghalayini, Mark D. Risser
Hugging Face Trending Papers
Aug 27

Why not to use the Gaussian kernel

The paper argues against using the Gaussian (squared exponential/RBF) kernel as a default in Gaussian process regression, citing its brittleness. It shows that the kernel leads to unrealistically small conditional variances, causing overconfidence in predictive uncertainty, and that this small variance induces numerical ill‑conditioning, necessitating tricks like nugget terms that alter the model. The authors attribute these issues to the kernel’s analytic, highly smooth nature and suggest that analytic stationary kernels in general should be avoided.

arXiv Machine Learning
Jun 17

Approximating Gaussian Whittle-Matern Fields over Well-Centered Triangulations of Riemannian Manifolds

arXiv:2606. 13827v2 Announce Type: replace-cross Abstract: Markovian Whittle-Mat\'ern fields have been convergently approximated by discrete Gauss Markov Random Fields (GMRFs) with sparse precision matrices using a Finite Element approximation of the two-parameter family, \[ (\kappa^2 - \Delta)^{\alpha/2} u = \mathcal{W}, \;\; \kappa \in \mathbb{R}, \; \alpha \in \mathbb{N}.

By Srinivas Nambirajan
arXiv Machine Learning
Sep 11

Sublinear Variational Optimization of Gaussian Mixture Models with Millions to Billions of Parameters

The paper introduces a highly efficient variational approximation for Gaussian Mixture Models (GMMs) with arbitrary covariances, integrated with mixtures of factor analyzers. This method reduces the per‑iteration runtime from ≠O(NCD^2) to a complexity that scales linearly with dimensionality D and sublinearly with the product NC. Experiments demonstrate sublinear scaling across the entire optimization, order‑of‑magnitude speed‑ups on large benchmarks, training of GMMs with over 10 billion parameters in under nine hours on a single CPU, and competitive zero‑shot image denoising performance.

By Sebastian Salwig, Till Kahlke, Florian Hirschberger, Dennis Forster, J\"org L\"ucke