arXiv Machine Learning

On Basis Function Selection for Sparse Gaussian Process Regression

The paper proposes three information‑theoretic criteria for selecting the most relevant basis functions in sparse Gaussian process regression, tailored to different levels of prior knowledge. Experiments on six UCI regression datasets and three basis families (HSGP, VFF, VISH) show that the no‑data criterion is a robust default, often outperforming simple truncation, while the data‑aware criteria yield significant improvements for HSGP. The study demonstrates that careful basis‑function selection can lead to better performance without increasing computational cost.

arXiv Machine Learning
Aug 24

Exact and general decoupled solutions of the LMC Multitask Gaussian Process model

The paper presents an exact, efficient solution for the Linear Model of Co‑regionalization (LMC) multitask Gaussian Process by decoupling latent processes under a mild noise‑model assumption. It introduces a full parametrization of the resulting projected LMC, enabling linear‑time optimization and simplifying tasks such as training updates and leave‑one‑out cross‑validation. Experiments on synthetic and real data demonstrate that projected LMC is competitive with state‑of‑the‑art multitask GP models while offering greater interpretability and computational ease.

By Olivier Truffinet (CEA Saclay), Karim Ammar (CEA Saclay), Jean-Philippe Argaud (EDF R&D), Bertrand Bouriquet (EDF)
arXiv Machine Learning
Aug 13

A Factor Graph Approach to Scalable Multi-Output Gaussian Process Regression

arXiv:2608. 11917v1 Announce Type: new Abstract: Multi-output Gaussian process regression scales cubically in the number of observations times outputs, and dense kernel-matrix methods need bespoke handling whenever different outputs are observed at different inputs.

By Wouter W. L. Nuijten, Esther G. van Pelt, Albert Podusenko, \.Ismail \c{S}en\"oz, Wouter M. Kouw
arXiv Machine Learning
Jul 27

gp2Scale: A Class of Compactly Supported Non-Stationary Kernels and Distributed Computing for Exact Gaussian Processes on 10 Million Data Points

arXiv:2512. 06143v2 Announce Type: replace Abstract: Despite a large corpus of recent work on scaling up Gaussian processes, a stubborn trade-off between computational speed, prediction and uncertainty quantification accuracy, and customizability persists.

By Marcus M. Noack, Mark D. Risser, Hengrui Luo, Vardaan Tekriwal, Ronald J. Pandolfi
arXiv AI
Sep 10

Revisiting Thinning Methods for Kernel Learning Problems

The paper introduces Backward Kernel Herding, an algorithm that iteratively removes data points to create representative subsets for kernel learning, achieving performance comparable to state‑of‑the‑art methods while speeding up subsampling when the reduced size is less than half the original dataset. It also proposes Flexible Kernel Thinning, an extension that allows construction of subsets of any size, not just successive halvings, and demonstrates that this method often yields the best predictive performance. Experiments on Gaussian Processes and Kernel Support Vector Machines show that Backward Kernel Herding excels in training‑time efficiency, while Flexible Kernel Thinning offers superior predictive accuracy and competitive memory usage, emphasizing the need to choose a reduction strategy based on the desired trade‑off between performance, cost, and memory.

By Blanca Cano-Camarero, Yago R. Aguado-Carrillo-de-Albornoz, \'Angela Fern\'andez-Pascual, Jos\'e R. Dorronsoro
arXiv Machine Learning
Sep 18

Online Adaptive Kernel Mixing for Gaussian Process Decision Making

The paper introduces HACK GPs, a method that treats kernel selection for Gaussian Processes as an online learning problem with expert advice. Each candidate kernel is viewed as a GP expert, and a distribution over these experts is updated online using AdaHedge based on a loss that reflects both function fit and task alignment. Two variants—Mixture of Gaussians and categorical sampling—are presented, with theoretical guarantees that the weight concentrates on the best kernel under a loss‑gap condition, and empirical results show robust performance across Bayesian optimization, level set estimation, and Bayesian active learning compared to standard kernels and simple ensembles.

By Kavin Aravindan, Mani Tej Sriram, Gautam Dasarathy, Tejas Bodas