arXiv:2501. 10538v3 Announce Type: replace Abstract: The practical success of deep learning has led to the discovery of several surprising phenomena.
By Ichiro Hashimoto, Stanislav Volgushev, Piotr Zwiernik
The paper analyzes training dynamics of multiclass logistic regression on high‑dimensional Gaussian mixture models with many classes. It finds that learning proceeds sequentially from the most to the least frequent classes and, when class priors follow a power‑law, the cross‑entropy risk evolves through an initial plateau, a power‑law decay phase, and a final convergence phase. The study also shows how model capacity and optimization trade‑off under a fixed compute budget, leading to a compute‑optimal scaling law that prescribes model size and training time as functions of compute.
By Konstantinos Christopher Tsiolis, Denny Wu, Christos Thrampoulidis, Murat A. Erdogdu
arXiv:2602. 12471v2 Announce Type: replace Abstract: We consider the optimization problem of minimizing the logistic loss with gradient descent to train a linear model for binary classification with separable data.
By Michael Crawshaw, Mingrui Liu
arXiv:2506. 06584v2 Announce Type: replace Abstract: Learning Gaussian Mixture Models (GMMs) is a fundamental problem in statistics and machine learning, with the Expectation-Maximization (EM) algorithm and its popular variant gradient EM being arguably the most widely used algorithms in practice.
By Mo Zhou, Weihang Xu, Maryam Fazel, Simon S. Du
arXiv:2606. 19876v1 Announce Type: new Abstract: The score matching problem is a central training objective in modern generative modeling, diffusion models, fitting unnormalized statistical models, and inverse problems.
By Alexander Tyurin
The paper proposes an analytic method for determining the optimal early‑stopping time in training neural networks, avoiding the need for gradient‑descent training. It uses Rademacher complexity with an L1‑norm to estimate generalization error, offering a more general approach than previous random‑matrix‑theory based methods. The framework is demonstrated on linear regression and extended to nonlinear neural networks via linear probing, as shown in a MNIST classification example.
By Duy Hoang, Bastien Berret, Olivier Bruneau, Laurent Fribourg