Universality of Benign Overfitting in Binary Linear Classification
arXiv:2501. 10538v3 Announce Type: replace Abstract: The practical success of deep learning has led to the discovery of several surprising phenomena.
arXiv:2608. 06250v1 Announce Type: cross Abstract: In overparameterised classification, training data can be linearly separable even when the underlying distribution is not.
arXiv:2501. 10538v3 Announce Type: replace Abstract: The practical success of deep learning has led to the discovery of several surprising phenomena.
The paper analyzes training dynamics of multiclass logistic regression on high‑dimensional Gaussian mixture models with many classes. It finds that learning proceeds sequentially from the most to the least frequent classes and, when class priors follow a power‑law, the cross‑entropy risk evolves through an initial plateau, a power‑law decay phase, and a final convergence phase. The study also shows how model capacity and optimization trade‑off under a fixed compute budget, leading to a compute‑optimal scaling law that prescribes model size and training time as functions of compute.
arXiv:2602. 12471v2 Announce Type: replace Abstract: We consider the optimization problem of minimizing the logistic loss with gradient descent to train a linear model for binary classification with separable data.
arXiv:2506. 06584v2 Announce Type: replace Abstract: Learning Gaussian Mixture Models (GMMs) is a fundamental problem in statistics and machine learning, with the Expectation-Maximization (EM) algorithm and its popular variant gradient EM being arguably the most widely used algorithms in practice.
arXiv:2606. 19876v1 Announce Type: new Abstract: The score matching problem is a central training objective in modern generative modeling, diffusion models, fitting unnormalized statistical models, and inverse problems.
The paper proposes an analytic method for determining the optimal early‑stopping time in training neural networks, avoiding the need for gradient‑descent training. It uses Rademacher complexity with an L1‑norm to estimate generalization error, offering a more general approach than previous random‑matrix‑theory based methods. The framework is demonstrated on linear regression and extended to nonlinear neural networks via linear probing, as shown in a MNIST classification example.
arXiv:2607. 21773v1 Announce Type: new Abstract: In this paper, we propose and study a robust variant of the smart predict-then-optimize approach that accounts for prediction shifts due to disturbance in the covariate feature space.
arXiv:2605. 02701v2 Announce Type: replace-cross Abstract: We propose a robust gradient estimator based on per-sample gradient clipping and analyze its properties both theoretically and empirically.
arXiv:2406. 04425v2 Announce Type: replace Abstract: A fundamental problem in machine learning is understanding the effect of early stopping on the parameters obtained and the generalization capabilities of the model.
arXiv:2606. 06469v1 Announce Type: cross Abstract: Let $S$ be the set of unit norm linear classifiers $\theta \in \mathbb{R}^d$ which correctly classify every point of a labeled dataset $(X_i,y_i)_{i=1}^n$, $X_i \in \mathbb{R}^d$, $y_i \in \{-1,+1\}$, with a possibly negative margin $\kappa$ fixed in advance.
arXiv:2604. 08625v2 Announce Type: replace-cross Abstract: We develop a theoretical framework for generalization in the interpolating regime of statistical learning.
Training neural networks requires balancing the trade-off between fitting the training data and achieving robust performance on unseen inputs. This ability, commonly referred to as generalizability, i...