arXiv Machine Learning

Benign Overfitting Does Not Occur in Diffusion Models

arXiv:2607. 02671v1 Announce Type: cross Abstract: Benign overfitting and double descent have come to shape our understanding of generalization in deep learning, establishing that overfitting is not only compatible with good generalization but can actively benefit it.

arXiv Machine Learning
Sep 23

Double Descent and Malign Overfitting in Diffusion Models

The paper investigates why diffusion models, unlike typical deep learning models, exhibit catastrophic overfitting when overparameterized. Through experiments on U‑Nets trained on CelebA and a random‑features theoretical analysis, it shows that the interpolation peak occurs at a model size proportional to the product of training samples and noise realizations, but the test loss starts to rise already at the number of samples, leading to memorization of the empirical score. Regularization techniques such as ridge penalties or early stopping can still make large models outperform smaller, unregularized ones.

By Rapha\"el Urfin, Tony Bonnaire, Giulio Biroli, Marc M\'ezard
arXiv Machine Learning
Aug 26

Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regime

The paper investigates diffusion models trained in a lazy high‑dimensional regime, extending benign overfitting theory to generative settings. By analyzing denoising score matching in a vector‑valued RKHS with an inner‑product kernel, the authors derive exact risk trajectories under gradient flow when the number of samples scales proportionally with dimensionality. These trajectories reveal three distinct phases—spectral generalization, noise‑dominated interpolation, and empirical Bayes memorization—whose interplay shapes the distribution of generated samples.

By Hugo Latourelle-Vigeant, Sinho Chewi, Aram-Alexandre Pooladian, John Sous, Theodor Misiakiewicz
arXiv Machine Learning
Jun 10

The Emergence of Reproducibility and Generalizability in Diffusion Models

arXiv:2310. 05264v5 Announce Type: replace Abstract: In this work, we investigate an intriguing and prevalent phenomenon of diffusion models which we term as "consistent model reproducibility": given the same starting noise input and a deterministic sampler, different diffusion models often yield remarkably similar outputs.

By Huijie Zhang, Jinfan Zhou, Yifu Lu, Minzhe Guo, Peng Wang, Liyue Shen, Qing Qu
arXiv Machine Learning
Sep 3

Perceptually Regularized Diffusion Model for Image Super-Resolution

The paper introduces a perceptually regularized diffusion framework for image super‑resolution, adding perceptual‑loss based regularization to the standard diffusion training objective. This approach incorporates prior knowledge to improve training convergence and encourages the recovery of meaningful image features. Experiments on benchmark datasets show enhanced perceptual quality while maintaining competitive distortion metrics.

By Chuxiangbo Wang, Pavithra Venkatachalapathy, Ying Liang, Min Wang, Jing Qin, Yifei Lou, Weihong Guo
arXiv Machine Learning
Sep 11

Smoothing the Score Function to Enhance Generalization in Diffusion Models

The paper investigates memorization in diffusion models, showing that the empirical score function is a weighted sum of Gaussian score functions with sharp softmax weights, causing individual training samples to dominate and lead to sampling collapse. By approximating this function with a neural network, the authors obtain a smoother representation that generalizes better. They introduce two techniques—Noise Unconditioning and Temperature Smoothing—to further reduce single‑sample dominance, and demonstrate improved generalization across multiple datasets while preserving generation quality.

By Xinyu Zhou, Jiawei Zhang, Stephen J. Wright
arXiv AI
Sep 10

DUA-D2C: Dynamic Uncertainty Aware Method for Overfitting Remediation in Deep Learning

The paper introduces DUA-D2C, a Dynamic Uncertainty-Aware Divide2Conquer method that improves overfitting remediation in deep learning. It refines the traditional Divide2Conquer approach by dynamically weighting subset models based on a composite score of accuracy and normalized prediction entropy, allowing the central model to learn more from generalizable and confident edge models. The authors provide theoretical justification, show reduced model variance, and demonstrate significant generalization gains across image, audio, and text benchmarks, even when combined with standard regularizers like Dropout.

By Md. Saiful Bari Siddiqui, Md Mohaiminul Islam, Md. Golam Rabiul Alam