How diffusion models circumvent the curse of dimensionality to learn complex distributions over high dimensional spaces from a finite training set, instead of memorizing it, remains a fundamental mystery. To address this, we introduce analytically tractable Bayesian information restricted diffusion (BIRD) models, in which each pixel observes restricted information about noisy data.
The paper investigates when language diffusion models, specifically Uniform-based Discrete Diffusion Models (UDDMs), shift from memorizing training data to generalizing to new data. It shows that UDDMs act as associative memories, forming basins of attraction around stored examples without requiring an explicit energy function. By measuring token recovery and conditional entropy, the authors identify a sharp transition governed by training set size, where memorization (vanishing entropy) gives way to generalization (finite entropy).
By Bao Pham, Mohammed J. Zaki, Luca Ambrogioni, Dmitry Krotov, Matteo Negri
arXiv:2310. 05264v5 Announce Type: replace Abstract: In this work, we investigate an intriguing and prevalent phenomenon of diffusion models which we term as "consistent model reproducibility": given the same starting noise input and a deterministic sampler, different diffusion models often yield remarkably similar outputs.
By Huijie Zhang, Jinfan Zhou, Yifu Lu, Minzhe Guo, Peng Wang, Liyue Shen, Qing Qu
arXiv:2605.12597v3 Announce Type: replace-cross
Abstract: Computational sampling has been central to the sciences since the mid-20th century. While machine-learning-based approaches have recently ena...
By Luca Maria Del Bono, Giulio Biroli, Patrick Charbonneau, Marylou Gabri\'e
The paper argues that diffusion models, originally developed for image synthesis, implicitly perform concept formation similar to the Cobweb cognitive model. Both models build hierarchical density structures using Gaussian prototypes, treat categorization as score‑following to reduce uncertainty, and exhibit a basic level of abstraction. The authors demonstrate this correspondence by extracting a diffusion hierarchy from MNIST and Fashion‑MNIST data and comparing its basic level to that of Cobweb, suggesting diffusion models can serve as a continuous, scalable instantiation of concept formation.
By Zekun Wang, Karthik Singaravadivelan, Christopher J. MacLellan
arXiv:2606. 09718v1 Announce Type: new Abstract: Diffusion models have demonstrated remarkable generative capabilities and have also emerged as powerful self-supervised representation learners, yet the connection between these two abilities remains less explored.
By Xiao Li, Yixuan Jia, Zekai Zhang, Xiang Li, Lianghe Shi, Jinxin Zhou, Zhihui Zhu, Liyue Shen, Qing Qu
The paper introduces a data‑free learning objective called relative trajectory balance for training diffusion models to sample from a posterior defined by a diffusion prior and an arbitrary black‑box constraint or likelihood. It proves asymptotic correctness of this objective and demonstrates its use across vision, language, and multimodal tasks, including classifier guidance, language infilling, and text‑to‑image generation. Additionally, the method is applied to continuous control with a score‑based behavior prior, achieving state‑of‑the‑art results in offline reinforcement learning.
By Siddarth Venkatraman, Moksh Jain, Luca Scimeca, Minsu Kim, Marcin Sendera, Mohsin Hasan, Luke Rowe, Sarthak Mittal, Pablo Lemos, Emmanuel Bengio, Alexandre Adam, Jarrid Rector-Brooks, Yoshua Bengio, Glen Berseth, Esmeralda S. Whitammer
arXiv:2602. 09651v2 Announce Type: replace-cross Abstract: Diffusion models do not recover semantic structure uniformly over time.
By Florian Handke, Dejan Stan\v{c}evi\'c, Felix Koulischer, Thomas Demeester, Luca Ambrogioni
arXiv:2602. 02908v2 Announce Type: replace-cross Abstract: Diffusion models trained on different, non-overlapping subsets of a dataset often produce strikingly similar outputs when given the same noise seed.
By Binxu Wang, Jacob Zavatone-Veth, Cengiz Pehlevan
arXiv:2606. 13796v1 Announce Type: cross Abstract: Recursive training of generative models on their own outputs can lead to model collapse, a compounding drift away from the true data distribution.
By Na\"il B. Khelifa, Richard E. Turner, Ramji Venkataramanan
The book "The Principles of Diffusion Models" outlines the foundational concepts behind diffusion models, tracing their evolution from a forward process that corrupts data into noise to a reverse process that reconstructs data. It presents three complementary perspectives—variational, score-based, and flow-based—each describing how a time-dependent velocity field transports a simple prior to the data distribution. The text also covers practical guidance for controllable generation, efficient solvers, and diffusion-inspired flow-map models, providing a mathematically grounded framework for readers with basic deep‑learning knowledge.
By Chieh-Hsin Lai, Yang Song, Dongjun Kim, Yuki Mitsufuji, Stefano Ermon
arXiv:2605. 13175v2 Announce Type: replace Abstract: Recent works have proposed incorporating heavy-tailed (HT) noise into diffusion- and flow-based generative models, with the goals of better recovering the tails of target distributions and improving generative diversity.
By Hamza Cherkaoui, H\'el\`ene Halconruy, Antonio Ocello