arXiv Machine Learning By Xinyu Zhou, Jiawei Zhang, Stephen J. Wright

Smoothing the Score Function to Enhance Generalization in Diffusion Models

Read the original on arXiv Machine Learning →

The paper investigates memorization in diffusion models, showing that the empirical score function is a weighted sum of Gaussian score functions with sharp softmax weights, causing individual training samples to dominate and lead to sampling collapse. By approximating this function with a neural network, the authors obtain a smoother representation that generalizes better. They introduce two techniques—Noise Unconditioning and Temperature Smoothing—to further reduce single‑sample dominance, and demonstrate improved generalization across multiple datasets while preserving generation quality.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 10

The Emergence of Reproducibility and Generalizability in Diffusion Models

arXiv:2310. 05264v5 Announce Type: replace Abstract: In this work, we investigate an intriguing and prevalent phenomenon of diffusion models which we term as "consistent model reproducibility": given the same starting noise input and a deterministic sampler, different diffusion models often yield remarkably similar outputs.

By Huijie Zhang, Jinfan Zhou, Yifu Lu, Minzhe Guo, Peng Wang, Liyue Shen, Qing Qu
arXiv Machine Learning
Jun 10

MAD: Manifold Attracted Diffusion

arXiv:2509. 24710v2 Announce Type: replace-cross Abstract: Score-based diffusion models are a highly effective method for generating samples from a distribution of images.

By Dennis Elbr\"achter, Giovanni S. Alberti, Matteo Santacesaria
arXiv Machine Learning
Aug 26

Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regime

The paper investigates diffusion models trained in a lazy high‑dimensional regime, extending benign overfitting theory to generative settings. By analyzing denoising score matching in a vector‑valued RKHS with an inner‑product kernel, the authors derive exact risk trajectories under gradient flow when the number of samples scales proportionally with dimensionality. These trajectories reveal three distinct phases—spectral generalization, noise‑dominated interpolation, and empirical Bayes memorization—whose interplay shapes the distribution of generated samples.

By Hugo Latourelle-Vigeant, Sinho Chewi, Aram-Alexandre Pooladian, John Sous, Theodor Misiakiewicz