Benign Overfitting in Linear Classifiers with a Bias Term
arXiv:2511. 12840v2 Announce Type: replace-cross Abstract: Overparameterized models often generalize well even when they interpolate noisy training data.
arXiv:2501. 10538v3 Announce Type: replace Abstract: The practical success of deep learning has led to the discovery of several surprising phenomena.
arXiv:2511. 12840v2 Announce Type: replace-cross Abstract: Overparameterized models often generalize well even when they interpolate noisy training data.
arXiv:2608. 06250v1 Announce Type: cross Abstract: In overparameterised classification, training data can be linearly separable even when the underlying distribution is not.
arXiv:2109.02355v2 Announce Type: replace Abstract: The last decade of progress in machine learning (ML), especially the deep learning era, has raised a number of scientific questions that challenge...
arXiv:2607. 02671v1 Announce Type: cross Abstract: Benign overfitting and double descent have come to shape our understanding of generalization in deep learning, establishing that overfitting is not only compatible with good generalization but can actively benefit it.
arXiv:2607. 13631v1 Announce Type: new Abstract: The Hessian matrix is an important quantity of interest when it comes to studying the loss landscape and optimization dynamics in deep learning, as well as designing measures of generalization, second-order learning algorithms, etc.
arXiv:2309. 15769v3 Announce Type: replace-cross Abstract: Recent advances in deep learning have highlighted the phenomenon of benign overfitting in overparameterized statistical models, sparking significant interest in understanding its foundations.
arXiv:2608. 08489v1 Announce Type: new Abstract: Neural network classifiers trained by cross-entropy minimization are highly sensitive to label noise and adversarial contamination.
arXiv:2205. 07739v4 Announce Type: replace-cross Abstract: Self-training (ST) is a simple yet effective semi-supervised learning method.
The Hessian matrix is an important quantity of interest when it comes to studying the loss landscape and optimization dynamics in deep learning, as well as designing measures of generalization, second-order learning algorithms, etc. Prior works have focused on empirical results or pursued a theoretical treatment under overly simplified settings.
arXiv:2601. 19791v4 Announce Type: replace Abstract: We study grokking, the onset of generalization long after overfitting, in a classical ridge regression setting.
arXiv:2606. 16050v1 Announce Type: cross Abstract: Robust deep learning under heavy-tailed and impulsive noise remains challenging because conventional losses such as mean squared error (MSE) exhibit unbounded sensitivity to outliers.
The paper investigates binary classification with abstention under separate class‑conditional error constraints, aiming to minimize abstention while keeping both error types below specified thresholds. It derives the distribution‑free minimax rate of excess abstention risk, introduces surrogate‑loss formulations for computational feasibility with models like neural networks, and provides finite‑sample guarantees for excess surrogate ambiguity risk. The authors also formulate the learning task as a constrained optimization problem, analyze its computational complexity in the convex setting, and empirically evaluate the approach against a competing method on several datasets.