Noise2Noise Revisited: Training Pair Distributions Dominate Loss Choice in Self-Supervised Denoising
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2608. 01032v1 Announce Type: new Abstract: Training error is what we can observe on a training set; test error is the quantity we actually care about.
The paper investigates why diffusion models, unlike typical deep learning models, exhibit catastrophic overfitting when overparameterized. Through experiments on U‑Nets trained on CelebA and a random‑features theoretical analysis, it shows that the interpolation peak occurs at a model size proportional to the product of training samples and noise realizations, but the test loss starts to rise already at the number of samples, leading to memorization of the empirical score. Regularization techniques such as ridge penalties or early stopping can still make large models outperform smaller, unregularized ones.
Conventional wisdom in deep learning holds that overparameterization---having more parameters $p$ than training samples $n$---is benign: larger models generalize better and, even without regularizatio...
arXiv:2608.23916v1 Announce Type: new Abstract: Denoising score matching trains diffusion models by regressing onto a conditional score, although generation ultimately requires the marginal score. Th...
arXiv:2610. 00436v1 Announce Type: new Abstract: Online batch selection fine-tunes a language model on the most useful part of each candidate batch.
arXiv:2607. 12360v1 Announce Type: new Abstract: The cooldown phase of a warmup-stable-decay (WSD) learning-rate schedule, now a default in large-model pretraining, lowers the final training loss in some settings and does nothing in others.