Simplex-to-Euclidean Bijection for Conjugate and Calibrated Multiclass Gaussian Process Classification
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2605. 10285v2 Announce Type: replace-cross Abstract: We present a theoretically grounded Gaussian process framework that leverages neural feature maps to construct expressive kernels.
The paper introduces a framework that learns the kernel used in kernel methods through alignment, leveraging the Collaborative Learning and Inference (CLaI) approach. It demonstrates that CLaI can be interpreted as a kernel alignment process and that its inference stage is equivalent to kernel Bayes classification with Parzen-window density estimation. By replacing cosine similarity with a learned Mahalanobis distance, the authors extend CLaI to multiclass classification, achieving higher accuracy, faster convergence, and lower calibration error on datasets such as CIFAR-10, PathMNIST, and SleepEDF, while also showing connections to Gaussian processes and competitive calibration in sepsis prediction.
arXiv:2607. 19498v1 Announce Type: cross Abstract: Gaussian process (GP) modeling is widely used in computational science and engineering.
arXiv:2607. 04977v1 Announce Type: new Abstract: Accurately estimating the unknown target label distribution is the critical first step for adapting to label shift.
arXiv:2609. 21085v1 Announce Type: cross Abstract: Gaussian processes (GPs) provide principled probabilistic predictions while encoding prior knowledge, including equivariances.
The paper introduces HACK GPs, a method that treats kernel selection for Gaussian Processes as an online learning problem with expert advice. Each candidate kernel is viewed as a GP expert, and a distribution over these experts is updated online using AdaHedge based on a loss that reflects both function fit and task alignment. Two variants—Mixture of Gaussians and categorical sampling—are presented, with theoretical guarantees that the weight concentrates on the best kernel under a loss‑gap condition, and empirical results show robust performance across Bayesian optimization, level set estimation, and Bayesian active learning compared to standard kernels and simple ensembles.