arXiv Machine Learning

A Variational Analysis of Kernel Learning with Learnable Linear Transformations

arXiv:2502. 11665v3 Announce Type: replace-cross Abstract: The classical kernel ridge regression problem aims to find the best fit for the output $Y$ as a function of the input data $X\in \mathbb{R}^d$, with a fixed choice of regularization term imposed by a given choice of a reproducing kernel Hilbert space, such as a Sobolev space.

arXiv Machine Learning
Sep 10

Learning Multi-Index Models with Hyper-Kernel Ridge Regression

arXiv:2510.02532v2 Announce Type: replace-cross Abstract: Deep neural networks excel in high-dimensional problems, outperforming models such as kernel methods, which suffer from the curse of dimensio...

By Shuo Huang, Hippolyte Labarri\`ere, Ernesto De Vito, Tomaso Poggio, Lorenzo Rosasco
arXiv Machine Learning
Sep 2

A Compositional Kernel Model for Feature Learning

The paper introduces a compositional variant of kernel ridge regression where the predictor reweights input coordinates, framing the approach as a variational problem to study feature learning in compositional architectures. It demonstrates that both global minimizers and stationary points can discard Gaussian noise variables while retaining relevant ones, and shows that α1-type kernels (e.g., Laplace) recover features contributing to nonlinear effects at stationary points, whereas Gaussian kernels recover only linear ones.

By Feng Ruan, Keli Liu, Michael Jordan
arXiv Machine Learning
Jun 9

Generalization in Nonlinear Least Squares via Learned Feature Geometry

arXiv:2606. 08799v1 Announce Type: cross Abstract: We study the generalization of ridge-regularized nonlinear least-squares models via on-average algorithmic stability, deriving error bounds for local minimizers in terms of a data-dependent effective dimension that reflects the geometry of the gradient model at the trained parameters, through the empirical Jacobian Gram matrix and a residual--curvature term.

By Ayub Kharel, Ilja Kuzborski, Patrick Rebeschini, Yasin Abbasi-Yadkori
arXiv Machine Learning
Aug 26

Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regime

The paper investigates diffusion models trained in a lazy high‑dimensional regime, extending benign overfitting theory to generative settings. By analyzing denoising score matching in a vector‑valued RKHS with an inner‑product kernel, the authors derive exact risk trajectories under gradient flow when the number of samples scales proportionally with dimensionality. These trajectories reveal three distinct phases—spectral generalization, noise‑dominated interpolation, and empirical Bayes memorization—whose interplay shapes the distribution of generated samples.

By Hugo Latourelle-Vigeant, Sinho Chewi, Aram-Alexandre Pooladian, John Sous, Theodor Misiakiewicz
arXiv Machine Learning
Jul 9

Fixed-Gaussian Spectral Algorithms: Minimax Optimal Rates for Misspecified Learning and Transfer

arXiv:2501. 10870v2 Announce Type: replace-cross Abstract: The principal objective of this work is twofold within nonparametric regression settings: (1) to establish the minimax optimal convergence rates for fixed-bandwidth Gaussian kernel spectral algorithms when the true regression function resides in a Sobolev space, and (2) to apply Gaussian spectral algorithms for achieving robust and adaptive transfer learning under concept shift.

By Haotian Lin, Matthew Reimherr
arXiv Machine Learning
Sep 16

A Weighted Kernel Method for Approximation that Adapts to Learned Multivariable Structure

The paper introduces Total Sensitivity Kernels (TSKs), a weighted ANOVA kernel framework that learns the importance of individual inputs and their interactions for approximating a multivariable black-box function from limited data. By selecting an RKHS where the target function has minimum norm, the authors derive a unique solution and prove consistency for finite-data interpolation. Numerical experiments show that adapting the kernel to the learned multivariable structure can significantly improve approximation accuracy compared to a standard product kernel.

By John E. Darges, Laura Weidensager
arXiv Machine Learning
Sep 17

Fast Learning Rates for Physics-Informed Kernel Methods

arXiv:2609. 18901v1 Announce Type: cross Abstract: In physics-informed machine learning, a target function $u^*$ is learned from noisy value observations $y_i=u^*(x_i)+ \varepsilon_i$, together with differential information, given either by noisy observations $d_j=(Du^*)(z_j)+\xi_j$ or by a known physical constraint $Du^*=v$.

By Luc Brogat-Motte, Joachim Bona-Pellissier, Giacomo Meanti, Lorenzo Rosasco
arXiv Machine Learning
Jun 2

Optimal Regularization for Performative Learning

arXiv:2510. 12249v2 Announce Type: replace Abstract: In performative learning, the data distribution reacts to the deployed model - for example, because strategic users adapt their features to game it - which creates a more complex dynamic than in classical supervised learning.

By Edwige Cyffers, Alireza Mirrokni, Marco Mondelli