The paper introduces Moment Guided Diffusion (MGD), a new method that blends diffusion-based generative modeling with classical maximum entropy techniques. MGD samples maximum entropy distributions by solving a stochastic differential equation that steers moments toward specified values in finite time, thereby avoiding the slow mixing of traditional MCMC or Langevin dynamics. The authors prove convergence to the maximum entropy distribution in the large-volatility limit and provide a tractable entropy estimator, demonstrating the method on financial time series, turbulent flows, and cosmological fields using wavelet scattering moments.
By Etienne Lempereur, Nathana\"el Cuvelle--Magar, Florentin Coeurdoux, St\'ephane Mallat, Eric Vanden-Eijnden
The paper addresses the challenge of guiding conditioned generative models that are diffusion processes with singular diffusion coefficients, where traditional conditional densities may be nonexistent or non‑smooth. It proposes using causal optimal transport to construct approximate loss functions that identify a minimum‑entropy control for guidance, relying on the predictable representation property of conditioned diffusion processes and well‑posed martingale problems à la Üstünel.
The paper introduces a new method for conditioning degenerate diffusion models, which are generative models that rely on diffusion processes with singular diffusion coefficients. Traditional approaches use score functions for guidance, but this work employs causal optimal transport to define approximate loss functions that can identify a minimum‑entropy control even when conditional densities are non‑existent or non‑smooth. The method hinges on the predictable representation property of conditioned diffusion processes and the well‑posedness of their martingale problem, following the framework of "Ust"unel.
By U\u{g}ur Ayd{\i}n, Tamer Ba\c{s}ar
arXiv:2409. 18804v3 Announce Type: replace-cross Abstract: Denoising Diffusion Probabilistic Models (DDPM) are powerful state-of-the-art methods used to generate synthetic data from high-dimensional data distributions and are widely used for image, audio, and video generation as well as many more applications in science and beyond.
By Iskander Azangulov, George Deligiannidis, Judith Rousseau
arXiv:2508. 03636v3 Announce Type: replace-cross Abstract: We propose a Likelihood Matching approach for training diffusion models by first establishing an equivalence between the likelihood of the target data distribution and a likelihood along the sample path of the reverse diffusion.
By Lei Qian, Wu Su, Yanqi Huang, Song Xi Chen
arXiv:2602. 09639v2 Announce Type: replace Abstract: Denoising diffusion models (DDMs) are state-of-the-art methods for learning densities from data across numerous domains, yet many aspects of the training and sampling pipeline remain poorly understood.
By Zahra Kadkhodaie, Aram-Alexandre Pooladian, Sinho Chewi, Eero Simoncelli
arXiv:2506. 00849v2 Announce Type: replace Abstract: Despite the empirical success of Diffusion Models (DMs) and Variational Autoencoders (VAEs), their generalization performance remains theoretically underexplored, especially lacking a full consideration of the shared encoder-generator structure.
By Qi Chen, Jierui Zhu, Florian Shkurti
arXiv:2609.00279v1 Announce Type: cross
Abstract: This work shows that diffusion models learned with standard denoising loss can provide effective global MCMC proposals for complex high-dimensional t...
By Mitch Hill
arXiv:2606. 04324v1 Announce Type: new Abstract: One of the primary challenges in Bayesian inference on the parameters of a diffusion model from discrete observations is the unavailability of an analytical expression for the transition density function between consecutive observation times, which is needed to derive the likelihood function.
By Riccardo Saporiti, Fabio Nobile
The paper revisits the continuous diffusion language model Plaid and introduces RePlaid, aligning its architecture with modern discrete diffusion models. RePlaid achieves a compute gap of only 20× compared to autoregressive models, surpasses Duo with fewer parameters, and outperforms MDLM in over‑trained settings. On OpenWebText, RePlaid sets a new state‑of‑the‑art continuous diffusion perplexity of 22.1 and demonstrates superior generation quality, while theoretical analysis links likelihood‑based training to linear cross‑entropy over time and structured embedding geometries.
By Zhihan Yang, Wei Guo, Shuibai Zhang, Subham Sekhar Sahoo, Yongxin Chen, Arash Vahdat, Morteza Mardani, John Thickstun
arXiv:2606. 18071v1 Announce Type: cross Abstract: Score-based diffusion models typically use Brownian perturbations, which provide tractable reverse-time dynamics but impose memoryless noising.
By Yusen Jia, Bingyan Han
The paper establishes a first‑order theoretical framework for diffusion models, showing that SDE‑based reverse‑time flows of both overdamped and underdamped Langevin diffusions contract relative Fisher divergences at explicit exponential rates when the stationary potential of the forward process is strongly convex. It further incorporates discretization to provide averaged first‑order stationarity bounds—sampling analogues of averaged gradient‑norm guarantees in nonconvex optimization—for samplers of both diffusion models. These results highlight a unique advantage of SDE‑based reverse diffusion over ODE‑based approaches, offering local convexity‑free certificates that ensure score consistency rather than global mode weights.
By Zhifeng Chen, Chenyang Jiang, Yazhen Wang