arXiv Machine Learning

Projection-based multifidelity linear regression for data-scarce applications

arXiv:2508. 08517v2 Announce Type: replace-cross Abstract: Surrogate modeling for systems with high-dimensional quantities of interest remains challenging, particularly when training data are costly to acquire.

arXiv Machine Learning
Jul 7

Efficient Cross-Validation for Sparse Linear Regression

arXiv:2306. 14851v5 Announce Type: replace-cross Abstract: Given a high-dimensional covariate matrix and a response vector, ridge-regularized sparse linear regression selects a subset of features that explains the relationship between covariates and the response in an interpretable manner.

By Ryan Cory-Wright, Andr\'es G\'omez
arXiv Machine Learning
Sep 1

Prediction-Powered Conditional Inference

arXiv:2603.05575v2 Announce Type: replace-cross Abstract: We study prediction-powered conditional inference in the setting where labeled data are scarce, unlabeled covariates are abundant, and a blac...

By Yang Sui, Jin Zhou, Hua Zhou, Xiaowu Dai
arXiv Machine Learning
Jun 3

LAMP: Data-Efficient Linear Affine Weight-Space Models for Parameter-Controlled 3D Shape Generation and Extrapolation

arXiv:2510. 22491v3 Announce Type: replace Abstract: Generating high-fidelity 3D geometries under explicit parameter constraints is central to engineering design, yet current methods often require large datasets and fail to provide reliable control beyond the training distribution.

By Ghadi Nehme, Yanxia Zhang, Dule Shu, Matt Klenk, Faez Ahmed
arXiv Machine Learning
Sep 10

Learning Multi-Index Models with Hyper-Kernel Ridge Regression

arXiv:2510.02532v2 Announce Type: replace-cross Abstract: Deep neural networks excel in high-dimensional problems, outperforming models such as kernel methods, which suffer from the curse of dimensio...

By Shuo Huang, Hippolyte Labarri\`ere, Ernesto De Vito, Tomaso Poggio, Lorenzo Rosasco
arXiv Machine Learning
Jul 7

Distribution-free Deviation Bounds and The Role of Domain Knowledge in Learning via Model Selection with Cross-validation Risk Estimation

arXiv:2303. 08777v3 Announce Type: replace-cross Abstract: Cross-validation is one of the most widely used tools for risk estimation and model selection in statistics and machine learning, yet its theoretical properties when embedded in a learning procedure remain insufficiently understood.

By Diego Marcondes, Cl\'audia Peixoto
arXiv Machine Learning
Sep 24

CORE-STACK+: Meta-Learning for Deep Stacked Generalization

CORE-STACK+ is a new meta‑learning framework for deep stacked generalization that tackles two key problems in heterogeneous vision ensembles: prediction‑space multicollinearity and calibration collapse. It introduces a four‑step preconditioning pipeline—kernelized redundancy filtering, a lightweight differentiable meta‑feature gate, a spectrum‑adaptive ridge penalty, and a Laplace‑approximate Bayesian blender—to jointly improve conditioning and calibration. Across six vision benchmarks, CORE‑STACK+ boosts accuracy, reduces model count and inference cost, and significantly lowers expected calibration error compared to existing methods.

By Noor Islam S. Mohammad