arXiv:2603. 22962v3 Announce Type: replace Abstract: We study the theoretical behavior of denoising score matching--the learning task associated to diffusion models--when the data distribution is supported on a low-dimensional manifold and the score is parameterized using a random feature neural network.
By Anand Jerry George, Nicolas Macris
arXiv:2606. 19894v1 Announce Type: new Abstract: The remarkable success of score-based diffusion models has spurred significant efforts to establish their theoretical foundations.
By Xinhe Mu, Zaijiu Shang, Zhaoqi Zhou, Chuan Zhou, Qi Meng, Guiying Yan, Zhiming Ma
The paper investigates diffusion models trained in a lazy high‑dimensional regime, extending benign overfitting theory to generative settings. By analyzing denoising score matching in a vector‑valued RKHS with an inner‑product kernel, the authors derive exact risk trajectories under gradient flow when the number of samples scales proportionally with dimensionality. These trajectories reveal three distinct phases—spectral generalization, noise‑dominated interpolation, and empirical Bayes memorization—whose interplay shapes the distribution of generated samples.
By Hugo Latourelle-Vigeant, Sinho Chewi, Aram-Alexandre Pooladian, John Sous, Theodor Misiakiewicz
arXiv:2606. 09705v1 Announce Type: new Abstract: Scientific generative modeling often requires size transfer, where models trained on small systems are evaluated on larger ones.
By Wenjie Xi
The paper introduces a new method for training data attribution in diffusion models called TID, which uses a local score discrepancy measure and can be estimated without retraining. It further distills this approach into TIDE, a forward‑only student that reproduces the teacher’s rankings using internal activations, achieving comparable accuracy at dramatically lower query cost. Experiments on CIFAR‑10, ArtBench‑10, and MS‑COCO show that TID outperforms existing methods and TIDE attributes samples in milliseconds, faster than generation itself.
By Shixuan Liu, Joan Serr\`a, Kin Wai Cheuk, Jinju Kim, Woosung Choi, Yukara Ikemiya, Wei-Hsiang Liao, Jiaqi W. Ma, Yuki Mitsufuji
The paper investigates memorization in diffusion models, showing that the empirical score function is a weighted sum of Gaussian score functions with sharp softmax weights, causing individual training samples to dominate and lead to sampling collapse. By approximating this function with a neural network, the authors obtain a smoother representation that generalizes better. They introduce two techniques—Noise Unconditioning and Temperature Smoothing—to further reduce single‑sample dominance, and demonstrate improved generalization across multiple datasets while preserving generation quality.
By Xinyu Zhou, Jiawei Zhang, Stephen J. Wright
arXiv:2410.11771v4 Announce Type: replace
Abstract: Many spatial models exhibit locality structures that effectively reduce their intrinsic dimensionality, enabling efficient approximation and sampli...
By Tiangang Cui, Shuigen Liu, Xin T. Tong
arXiv:2310. 05264v5 Announce Type: replace Abstract: In this work, we investigate an intriguing and prevalent phenomenon of diffusion models which we term as "consistent model reproducibility": given the same starting noise input and a deterministic sampler, different diffusion models often yield remarkably similar outputs.
By Huijie Zhang, Jinfan Zhou, Yifu Lu, Minzhe Guo, Peng Wang, Liyue Shen, Qing Qu
arXiv:2605.12597v3 Announce Type: replace-cross
Abstract: Computational sampling has been central to the sciences since the mid-20th century. While machine-learning-based approaches have recently ena...
By Luca Maria Del Bono, Giulio Biroli, Patrick Charbonneau, Marylou Gabri\'e
arXiv:2509. 24710v2 Announce Type: replace-cross Abstract: Score-based diffusion models are a highly effective method for generating samples from a distribution of images.
By Dennis Elbr\"achter, Giovanni S. Alberti, Matteo Santacesaria
arXiv:2409. 02426v5 Announce Type: replace Abstract: Despite their empirical success across a wide range of generative tasks, the fundamental principles underlying the ability of diffusion models to learn data distributions are poorly understood.
By Peng Wang, Huijie Zhang, Zekai Zhang, Siyi Chen, Yi Ma, Qing Qu
arXiv:2501. 12982v3 Announce Type: replace-cross Abstract: This paper investigates how diffusion generative models leverage (unknown) low-dimensional structure to accelerate sampling.
By Jiadong Liang, Zhihan Huang, Yuxin Chen