arXiv Machine Learning

Diffusion Models Observe Only Gradients: A Geometric Perspective on Score Matching Errors

arXiv:2606. 06179v1 Announce Type: cross Abstract: Score-based diffusion models are typically trained by minimizing the $L^2$ score matching error, and standard theoretical analyses rely on this quantity to bound the sampling discrepancy between the learned and target distributions.

arXiv Machine Learning
Jul 20

Diffusion models recover accurate mixture weights despite score function insensitivity

arXiv:2607. 15485v1 Announce Type: new Abstract: Score-based generative models exhibit a puzzling behavior: they often appear to cover all modes of a target multimodal distribution and yet may fail to learn the correct relative mode amplitudes, which can be interpreted as mixture weights.

By Andrew Dennehy, Ramchandran Muthukumar, Rebecca Willett, Nisha Chandramoorthy
arXiv Machine Learning
Jul 17

The Effect of Stochasticity in Score-Based Diffusion Sampling: a KL Divergence Analysis

arXiv:2506. 11378v3 Announce Type: replace Abstract: Sampling in score-based diffusion models can be performed by solving either a reverse-time stochastic differential equation (SDE) parameterized by an arbitrary stochasticity function or a probability flow ODE, corresponding to setting this stochasticity function to zero.

By Bernardo P. Schaeffer, Ricardo M. S. Rosa, Glauco Valle
arXiv Machine Learning
Jun 19

Fisher-Geometric Sharpness and the Implicit Bias of SGD toward Flat Minima

arXiv:2606. 20469v1 Announce Type: new Abstract: A widely held intuition in deep learning is that stochastic gradient descent (SGD) implicitly favors flat minima and that flat minima generalize better, but standard Euclidean measures of flatness such as the trace or maximum eigenvalue of the loss Hessian are not invariant under reparametrizations that preserve the network function, which undermines the theoretical foundations of this narrative.

By Md Sakir Ahmed, Kumaresh Sarmah, Hemen Dutta