arXiv Machine Learning By Rapha\"el Urfin, Tony Bonnaire, Giulio Biroli, Marc M\'ezard

Double Descent and Malign Overfitting in Diffusion Models

Read the original on arXiv Machine Learning →

The paper investigates why diffusion models, unlike typical deep learning models, exhibit catastrophic overfitting when overparameterized. Through experiments on U‑Nets trained on CelebA and a random‑features theoretical analysis, it shows that the interpolation peak occurs at a model size proportional to the product of training samples and noise realizations, but the test loss starts to rise already at the number of samples, leading to memorization of the empirical score. Regularization techniques such as ridge penalties or early stopping can still make large models outperform smaller, unregularized ones.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 7

Benign Overfitting Does Not Occur in Diffusion Models

arXiv:2607. 02671v1 Announce Type: cross Abstract: Benign overfitting and double descent have come to shape our understanding of generalization in deep learning, establishing that overfitting is not only compatible with good generalization but can actively benefit it.

By Tyler Farghly, Benjamin Dupuis, Alain Durmus, Umut Simsekli