Benign interpolation and Occam's razor
arXiv:2608. 03386v1 Announce Type: new Abstract: Contemporary deep learning methods generalize well even when they fit their training data perfectly, a phenomenon known as benign interpolation.
arXiv:2605. 29819v2 Announce Type: replace Abstract: This work investigates theoretically the interplay between interpolation and aggregation in regression.
arXiv:2608. 03386v1 Announce Type: new Abstract: Contemporary deep learning methods generalize well even when they fit their training data perfectly, a phenomenon known as benign interpolation.
arXiv:2606. 24418v1 Announce Type: new Abstract: Data augmentation is a simple and model-agnostic approach for exploiting known invariances in learning problems.
arXiv:2511.22435v2 Announce Type: replace Abstract: Invariant learning on graphs aims to build predictors that rely on causal substructures rather than on environment-specific shortcuts. Current meth...
The paper argues that typical tabular machine learning benchmarks, which aggregate results by averaging scores or ranks, can hide which models are essential for achieving the best performance on specific datasets. It proposes evaluating models against a data‑centric peak performance frontier, classifying them as irreplaceable, sufficient, redundant, or fallible based on their position relative to other models. Applying this to the TabArena benchmark shows that common aggregation metrics mainly capture consistency and failure avoidance, but fail to reflect dataset‑specific strengths, leading to a misalignment between aggregate rewards and true model utility.
The paper investigates regression with bounded responses, comparing two learning frameworks: model selection aggregation, which requires improper algorithms to achieve minimax excess risk, and universal learning, where empirical risk minimization suffices for exponential learning rates. For finite hypothesis classes, the authors show that the $Q$-aggregation estimator simultaneously attains minimax optimal tails and exponential universal rates, while other common estimators fail to do so. For countably infinite classes, they prove an inherent trade‑off between exponential universal and minimax uniform rates, resolved by combining optimal algorithms from each framework via $Q$-aggregation.
arXiv:2501. 18530v3 Announce Type: replace-cross Abstract: We consider a teacher-student model of supervised learning with a fully-trained two-layer neural network whose width $k$ and input dimension $d$ are large and proportional.
Tabular machine learning benchmarks typically summarize performance by averaging scores, ranks, or pairwise wins across datasets. Such aggregates are useful for selecting robust default models, but th...
arXiv:2607. 24732v1 Announce Type: cross Abstract: Motivated by learning from heterogeneous and overlapping data providers, we study a stylized model of distribution learning from restricted conditional samples.
arXiv:2607. 07680v1 Announce Type: cross Abstract: Many machine learning models are defined for inputs of different sizes, such as point clouds containing different numbers of points, sequences of tokens of different lengths, and graphs on different numbers of nodes.
arXiv:2606. 28123v1 Announce Type: new Abstract: Last-iterate convergence and generalization guarantees in first-order convex learning hinge on the monotonicity of the update operator.
arXiv:2607. 08170v1 Announce Type: new Abstract: Zero-shot model size interpolation aims to create new models of intermediate target sizes by combining existing models without additional training.
arXiv:2603. 01346v2 Announce Type: replace Abstract: We revisit the framework of Smart PAC learning, which seeks supervised learners which compete with semi-supervised learners that are provided full knowledge of the marginal distribution on unlabeled data.