The Fourth Quadrant: A Stylized View of Benign Misfitting
arXiv:2608. 01032v1 Announce Type: new Abstract: Training error is what we can observe on a training set; test error is the quantity we actually care about.
arXiv:2608. 01032v1 Announce Type: new Abstract: Training error is what we can observe on a training set; test error is the quantity we actually care about.
The paper investigates why diffusion models, unlike typical deep learning models, exhibit catastrophic overfitting when overparameterized. Through experiments on U‑Nets trained on CelebA and a random‑features theoretical analysis, it shows that the interpolation peak occurs at a model size proportional to the product of training samples and noise realizations, but the test loss starts to rise already at the number of samples, leading to memorization of the empirical score. Regularization techniques such as ridge penalties or early stopping can still make large models outperform smaller, unregularized ones.
Conventional wisdom in deep learning holds that overparameterization---having more parameters $p$ than training samples $n$---is benign: larger models generalize better and, even without regularizatio...
arXiv:2608.23916v1 Announce Type: new Abstract: Denoising score matching trains diffusion models by regressing onto a conditional score, although generation ultimately requires the marginal score. Th...
arXiv:2610. 00436v1 Announce Type: new Abstract: Online batch selection fine-tunes a language model on the most useful part of each candidate batch.
arXiv:2607. 12360v1 Announce Type: new Abstract: The cooldown phase of a warmup-stable-decay (WSD) learning-rate schedule, now a default in large-model pretraining, lowers the final training loss in some settings and does nothing in others.
arXiv:2604. 21395v3 Announce Type: replace-cross Abstract: Ordinary supervised training minimises the task loss and then stops.
arXiv:2605.08144v2 Announce Type: replace-cross Abstract: Training a diffusion model involves two sources of randomness for each data sample: the timestep and the Gaussian noise realization. The time...
arXiv:2606. 16050v1 Announce Type: cross Abstract: Robust deep learning under heavy-tailed and impulsive noise remains challenging because conventional losses such as mean squared error (MSE) exhibit unbounded sensitivity to outliers.
arXiv:2410.18321v3 Announce Type: replace Abstract: Confidence calibration matters wherever a classifier's probabilities, not just its labels, are consumed downstream. We study Focal Calibration Loss...
arXiv:2609.15825v1 Announce Type: new Abstract: A self-supervised encoder is trained once, frozen, and reused through lightweight probes on tasks nobody named at training time; the practitioner's que...
arXiv:2609.26272v1 Announce Type: new Abstract: Neural samplers are trained against an unnormalised target $\tilde\pi=e^{-E}$ with no samples from $\pi$, which leaves the practitioner with no way to...