arXiv:2502. 11665v3 Announce Type: replace-cross Abstract: The classical kernel ridge regression problem aims to find the best fit for the output $Y$ as a function of the input data $X\in \mathbb{R}^d$, with a fixed choice of regularization term imposed by a given choice of a reproducing kernel Hilbert space, such as a Sobolev space.
By Yang Li, Feng Ruan
arXiv:2607. 00257v1 Announce Type: new Abstract: Accurate prediction of complex dynamical systems from noisy measurements remains a significant challenge in scientific computing.
By Max Kreider, John Harlim, Daning Huang
arXiv:2510.02532v2 Announce Type: replace-cross
Abstract: Deep neural networks excel in high-dimensional problems, outperforming models such as kernel methods, which suffer from the curse of dimensio...
By Shuo Huang, Hippolyte Labarri\`ere, Ernesto De Vito, Tomaso Poggio, Lorenzo Rosasco
arXiv:2604. 00316v2 Announce Type: replace-cross Abstract: Grokking occurs when a model achieves high training accuracy but generalization to unseen test points happens long after that.
By Marcel Tom\`as Bernal, Neil Rohit Mallinar, Mikhail Belkin
arXiv:2607.28844v2 Announce Type: replace-cross
Abstract: Bayesian Additive Regression Trees (BART) have shown state-of-the-art performance in both prediction and causal inference problems. Previous...
By Cory McCartan, Melody Huang
arXiv:2606. 00512v1 Announce Type: new Abstract: In many modern machine learning pipelines, abundant pretrained representations serve as noisy proxy covariates, while task-specific labels remain scarce.
By Kwangho Kim, Jisu Kim
arXiv:2606. 08799v1 Announce Type: cross Abstract: We study the generalization of ridge-regularized nonlinear least-squares models via on-average algorithmic stability, deriving error bounds for local minimizers in terms of a data-dependent effective dimension that reflects the geometry of the gradient model at the trained parameters, through the empirical Jacobian Gram matrix and a residual--curvature term.
By Ayub Kharel, Ilja Kuzborski, Patrick Rebeschini, Yasin Abbasi-Yadkori
The paper studies the numerical solution of the Beurling‑LASSO (BLASSO) for estimating Gaussian mixture models (GMMs) with unknown numbers of components and unknown diagonal covariance matrices. It introduces a Conic Particle Gradient Descent (CPGD) algorithm that incorporates Riemannian gradient descent to respect the Fisher‑Rao geometry of Gaussian distributions. The authors provide theoretical convergence guarantees, including exponential local convergence under a non‑degeneracy condition related to component separation, and demonstrate through numerical experiments that CPGD is more robust to overspecification of components than the EM algorithm.
By Romane Giard, Yohann De Castro, Roland Denis, Cl\'ement Marteau
arXiv:2608. 11831v1 Announce Type: new Abstract: Learning mappings between infinite-dimensional objects is a central challenge in scientific machine learning.
By Adrien Weihs, Chunyang Liao, Jingmin Sun, Hayden Schaeffer
arXiv:2206. 08598v2 Announce Type: replace Abstract: A common way to analyze learning of statistical models is to consider operations in the models parameter space, however this becomes challenging when there is no one-to-one mapping between the parameter space and the underlying statistical model space.
By Pascal Mattia Esser, Frank Nielsen
arXiv:2606. 21199v2 Announce Type: replace-cross Abstract: We introduce a semi-parametric framework for nonlinear system identification, which decouples discrepancy functions from physics-based components.
By Swapnil Manna, Timothy J. Rogers, Lawrence Bull
arXiv:2505.09716v3 Announce Type: replace-cross
Abstract: Out-of-distribution (OOD) generalisation is considered a hallmark of human and animal intelligence. To achieve OOD through composition, a sys...
By George Dimitriadis, Spyridon Samothrakis