arXiv Machine Learning

Distribution-Free Uncertainty Quantification for Kernel Methods by Gradient Perturbations

arXiv Machine Learning
Sep 17

Fast Learning Rates for Physics-Informed Kernel Methods

arXiv:2609. 18901v1 Announce Type: cross Abstract: In physics-informed machine learning, a target function $u^*$ is learned from noisy value observations $y_i=u^*(x_i)+ \varepsilon_i$, together with differential information, given either by noisy observations $d_j=(Du^*)(z_j)+\xi_j$ or by a known physical constraint $Du^*=v$.

By Luc Brogat-Motte, Joachim Bona-Pellissier, Giacomo Meanti, Lorenzo Rosasco
arXiv Machine Learning
Jul 28

Minimax Lower Bounds of Kernel Discrepancy Estimation: MMD, HSIC, KSD

arXiv:2607. 24235v1 Announce Type: cross Abstract: Over the past 20 years, kernel discrepancies have been leveraged as a highly powerful tool for quantifying the disagreement of distributions, with numerous successful applications in two-sample, goodness-of-fit, and independence testing, among others.

By Jose Cribeiro-Ramallo, Florian Kalinke, Zolt\'an Szab\'o
arXiv Machine Learning
Sep 17

Efficient Robust Learning at the Information-Theoretic Limit

The paper presents a polynomial‑time algorithm for robustly learning Boolean concept classes with respect to a fixed distribution, achieving the optimal error rate of η + ε where η is the noise rate. It builds on Blanc’s earlier, computationally inefficient algorithm and introduces no‑regret learners to overcome the previous limitations. Additionally, the authors provide an efficient method that does not require an ERM oracle for any function class admitting sandwiching polynomials under hypercontractive distributions, including a first polynomial‑time solution for learning halfspaces with Gaussian marginals at error η + ε.

By Adam R. Klivans, Konstantinos Stavropoulos, Sergei Tikhonov, Arsen Vasilyan
Hugging Face Trending Papers
Aug 27

Why not to use the Gaussian kernel

The paper argues against using the Gaussian (squared exponential/RBF) kernel as a default in Gaussian process regression, citing its brittleness. It shows that the kernel leads to unrealistically small conditional variances, causing overconfidence in predictive uncertainty, and that this small variance induces numerical ill‑conditioning, necessitating tricks like nugget terms that alter the model. The authors attribute these issues to the kernel’s analytic, highly smooth nature and suggest that analytic stationary kernels in general should be avoided.

arXiv Statistics ML
Sep 10

MiNCE: Nonparametric, Strongly Consistent Confidence Envelopes for Band-Limited Functions and their Smoothed Spectra

The paper introduces MiNCE, a nonparametric framework for constructing minimum‑norm confidence envelopes that provide nonasymptotic, simultaneous confidence regions for band‑limited functions using Reproducing Kernel Hilbert Spaces. It proves strong uniform consistency of these envelopes for both noise‑free and noisy data under mild noise assumptions, and extends the results to the frequency domain to yield consistent confidence bands for smoothed spectra. Numerical experiments in nonparametric regression and spectral estimation confirm the theoretical findings, showing the envelopes contract toward the target function as sample size grows.

By Bal\'azs Csan\'ad Cs\'aji, B\'alint Horv\'ath
arXiv Machine Learning
4d ago

Learning Distributionally Robust First-Order Methods for Convex Optimization

The paper introduces a distributionally robust method for learning hyperparameters of first‑order convex optimization algorithms. By minimizing a Wasserstein‑robust performance estimation problem over a dataset of problem instances, the approach interpolates between classical learning‑to‑optimize (L2O) and worst‑case PEP design. The authors solve the resulting problem with stochastic gradient descent, provide high‑probability risk bounds, and demonstrate that the learned algorithms outperform both worst‑case optimal and vanilla L2O baselines on logistic regression, LASSO, and linear programming tasks.

By Vinit Ranjan, Jisun Park, Bartolomeo Stellato
arXiv Machine Learning
Jul 27

gp2Scale: A Class of Compactly Supported Non-Stationary Kernels and Distributed Computing for Exact Gaussian Processes on 10 Million Data Points

arXiv:2512. 06143v2 Announce Type: replace Abstract: Despite a large corpus of recent work on scaling up Gaussian processes, a stubborn trade-off between computational speed, prediction and uncertainty quantification accuracy, and customizability persists.

By Marcus M. Noack, Mark D. Risser, Hengrui Luo, Vardaan Tekriwal, Ronald J. Pandolfi