arXiv:2502. 11665v3 Announce Type: replace-cross Abstract: The classical kernel ridge regression problem aims to find the best fit for the output $Y$ as a function of the input data $X\in \mathbb{R}^d$, with a fixed choice of regularization term imposed by a given choice of a reproducing kernel Hilbert space, such as a Sobolev space.
By Yang Li, Feng Ruan
arXiv:2601. 19791v4 Announce Type: replace Abstract: We study grokking, the onset of generalization long after overfitting, in a classical ridge regression setting.
By Mingyue Xu, Gal Vardi, Itay Safran
arXiv:2607. 00257v1 Announce Type: new Abstract: Accurate prediction of complex dynamical systems from noisy measurements remains a significant challenge in scientific computing.
By Max Kreider, John Harlim, Daning Huang
arXiv:2506. 13139v3 Announce Type: replace-cross Abstract: Modern Machine Learning (ML) and Deep Neural Networks (DNNs) often operate on high-dimensional data and rely on overparameterized models, where classical low-dimensional intuitions break down.
By Zhenyu Liao, Michael W. Mahoney
arXiv:2609.38842v1 Announce Type: cross
Abstract: Learn-then-differentiate (LTD) estimates gradients by fitting a model to simulation outputs and differentiating it. We develop a unified framework ex...
By Nifei Lin, Qingkai Zhang, L. Jeff Hong
arXiv:2608. 11831v1 Announce Type: new Abstract: Learning mappings between infinite-dimensional objects is a central challenge in scientific machine learning.
By Adrien Weihs, Chunyang Liao, Jingmin Sun, Hayden Schaeffer
arXiv:2606. 06772v1 Announce Type: cross Abstract: Understanding the generalization performance of over-parameterized neural networks has become a central topic in deep learning theory.
By Junyu Zhou, Puyu Wang, Yunwen Lei, Marius Kloft, Yiming Ying
The paper introduces a compositional variant of kernel ridge regression where the predictor reweights input coordinates, framing the approach as a variational problem to study feature learning in compositional architectures. It demonstrates that both global minimizers and stationary points can discard Gaussian noise variables while retaining relevant ones, and shows that α1-type kernels (e.g., Laplace) recover features contributing to nonlinear effects at stationary points, whereas Gaussian kernels recover only linear ones.
By Feng Ruan, Keli Liu, Michael Jordan
arXiv:2602. 02431v2 Announce Type: replace-cross Abstract: It is folklore that reusing training data more than once can improve the statistical efficiency of gradient-based learning.
By Filip Kova\v{c}evi\'c, Hong Chang Ji, Denny Wu, Mahdi Soltanolkotabi, Marco Mondelli
arXiv:2508. 08517v2 Announce Type: replace-cross Abstract: Surrogate modeling for systems with high-dimensional quantities of interest remains challenging, particularly when training data are costly to acquire.
By Vignesh Sella, Julie Pham, Karen Willcox, Anirban Chaudhuri
arXiv:2406. 04425v2 Announce Type: replace Abstract: A fundamental problem in machine learning is understanding the effect of early stopping on the parameters obtained and the generalization capabilities of the model.
By Rishi Sonthalia, Jackie Lok, Elizaveta Rebrova
arXiv:2609.38011v1 Announce Type: new
Abstract: Modern machine learning systems are trained on mixtures of data from different domains, and choosing the right mixture can substantially improve downst...
By Diyuan Wu, Lehan Chen, Theodor Misiakiewicz, Marco Mondelli