arXiv Machine Learning

A Compositional Kernel Model for Feature Learning

The paper introduces a compositional variant of kernel ridge regression where the predictor reweights input coordinates, framing the approach as a variational problem to study feature learning in compositional architectures. It demonstrates that both global minimizers and stationary points can discard Gaussian noise variables while retaining relevant ones, and shows that α1-type kernels (e.g., Laplace) recover features contributing to nonlinear effects at stationary points, whereas Gaussian kernels recover only linear ones.

arXiv Machine Learning
Aug 13

A Variational Analysis of Kernel Learning with Learnable Linear Transformations

arXiv:2502. 11665v3 Announce Type: replace-cross Abstract: The classical kernel ridge regression problem aims to find the best fit for the output $Y$ as a function of the input data $X\in \mathbb{R}^d$, with a fixed choice of regularization term imposed by a given choice of a reproducing kernel Hilbert space, such as a Sobolev space.

By Yang Li, Feng Ruan
arXiv Machine Learning
Sep 10

Learning Multi-Index Models with Hyper-Kernel Ridge Regression

arXiv:2510.02532v2 Announce Type: replace-cross Abstract: Deep neural networks excel in high-dimensional problems, outperforming models such as kernel methods, which suffer from the curse of dimensio...

By Shuo Huang, Hippolyte Labarri\`ere, Ernesto De Vito, Tomaso Poggio, Lorenzo Rosasco
arXiv Machine Learning
Jun 9

Generalization in Nonlinear Least Squares via Learned Feature Geometry

arXiv:2606. 08799v1 Announce Type: cross Abstract: We study the generalization of ridge-regularized nonlinear least-squares models via on-average algorithmic stability, deriving error bounds for local minimizers in terms of a data-dependent effective dimension that reflects the geometry of the gradient model at the trained parameters, through the empirical Jacobian Gram matrix and a residual--curvature term.

By Ayub Kharel, Ilja Kuzborski, Patrick Rebeschini, Yasin Abbasi-Yadkori
arXiv Statistics ML
Sep 25

Riemannian Gradient Descent for Gaussian Mixture Models with unknown diagonal covariances

The paper studies the numerical solution of the Beurling‑LASSO (BLASSO) for estimating Gaussian mixture models (GMMs) with unknown numbers of components and unknown diagonal covariance matrices. It introduces a Conic Particle Gradient Descent (CPGD) algorithm that incorporates Riemannian gradient descent to respect the Fisher‑Rao geometry of Gaussian distributions. The authors provide theoretical convergence guarantees, including exponential local convergence under a non‑degeneracy condition related to component separation, and demonstrate through numerical experiments that CPGD is more robust to overspecification of components than the EM algorithm.

By Romane Giard, Yohann De Castro, Roland Denis, Cl\'ement Marteau