arXiv Machine Learning

Conservation Laws for Diffusion Models

arXiv:2607. 10067v1 Announce Type: new Abstract: While autoregressive models optimize the exact data likelihood via the chain rule, diffusion models are typically trained with denoising objectives.

arXiv Machine Learning
Sep 10

MGD: Moment Guided Diffusion for Maximum Entropy Generation

The paper introduces Moment Guided Diffusion (MGD), a new method that blends diffusion-based generative modeling with classical maximum entropy techniques. MGD samples maximum entropy distributions by solving a stochastic differential equation that steers moments toward specified values in finite time, thereby avoiding the slow mixing of traditional MCMC or Langevin dynamics. The authors prove convergence to the maximum entropy distribution in the large-volatility limit and provide a tractable entropy estimator, demonstrating the method on financial time series, turbulent flows, and cosmological fields using wavelet scattering moments.

By Etienne Lempereur, Nathana\"el Cuvelle--Magar, Florentin Coeurdoux, St\'ephane Mallat, Eric Vanden-Eijnden
Hugging Face Trending Papers
Sep 3

Conditioning Degenerate Diffusion Models

The paper addresses the challenge of guiding conditioned generative models that are diffusion processes with singular diffusion coefficients, where traditional conditional densities may be nonexistent or non‑smooth. It proposes using causal optimal transport to construct approximate loss functions that identify a minimum‑entropy control for guidance, relying on the predictable representation property of conditioned diffusion processes and well‑posed martingale problems à la Üstünel.

arXiv Machine Learning
Sep 4

Conditioning Degenerate Diffusion Models

The paper introduces a new method for conditioning degenerate diffusion models, which are generative models that rely on diffusion processes with singular diffusion coefficients. Traditional approaches use score functions for guidance, but this work employs causal optimal transport to define approximate loss functions that can identify a minimum‑entropy control even when conditional densities are non‑existent or non‑smooth. The method hinges on the predictable representation property of conditioned diffusion processes and the well‑posedness of their martingale problem, following the framework of "Ust"unel.

By U\u{g}ur Ayd{\i}n, Tamer Ba\c{s}ar
arXiv Machine Learning
Aug 10

Convergence of Diffusion Models Under the Manifold Hypothesis in High-Dimensions

arXiv:2409. 18804v3 Announce Type: replace-cross Abstract: Denoising Diffusion Probabilistic Models (DDPM) are powerful state-of-the-art methods used to generate synthetic data from high-dimensional data distributions and are widely used for image, audio, and video generation as well as many more applications in science and beyond.

By Iskander Azangulov, George Deligiannidis, Judith Rousseau
arXiv Machine Learning
Jul 14

Likelihood Matching for Diffusion Models

arXiv:2508. 03636v3 Announce Type: replace-cross Abstract: We propose a Likelihood Matching approach for training diffusion models by first establishing an equivalence between the likelihood of the target data distribution and a likelihood along the sample path of the reverse diffusion.

By Lei Qian, Wu Su, Yanqi Huang, Song Xi Chen
arXiv Machine Learning
Jun 4

Neural Galerkin Normalizing Flows for Bayesian Inference of Diffusions with Inaccessible Boundaries

arXiv:2606. 04324v1 Announce Type: new Abstract: One of the primary challenges in Bayesian inference on the parameters of a diffusion model from discrete observations is the unavailability of an analytical expression for the transition density function between consecutive observation times, which is needed to derive the likelihood function.

By Riccardo Saporiti, Fabio Nobile
arXiv Machine Learning
Sep 11

Continuous Diffusion Scales Competitively with Discrete Diffusion for Language

The paper revisits the continuous diffusion language model Plaid and introduces RePlaid, aligning its architecture with modern discrete diffusion models. RePlaid achieves a compute gap of only 20× compared to autoregressive models, surpasses Duo with fewer parameters, and outperforms MDLM in over‑trained settings. On OpenWebText, RePlaid sets a new state‑of‑the‑art continuous diffusion perplexity of 22.1 and demonstrates superior generation quality, while theoretical analysis links likelihood‑based training to linear cross‑entropy over time and structured embedding geometries.

By Zhihan Yang, Wei Guo, Shuibai Zhang, Subham Sekhar Sahoo, Yongxin Chen, Arash Vahdat, Morteza Mardani, John Thickstun
arXiv AI
Jun 17

Volterra Generative Models

arXiv:2606. 18071v1 Announce Type: cross Abstract: Score-based diffusion models typically use Brownian perturbations, which provide tractable reverse-time dynamics but impose memoryless noising.

By Yusen Jia, Bingyan Han
arXiv Statistics ML
6d ago

First-Order Stationarity of Reverse Diffusions

The paper establishes a first‑order theoretical framework for diffusion models, showing that SDE‑based reverse‑time flows of both overdamped and underdamped Langevin diffusions contract relative Fisher divergences at explicit exponential rates when the stationary potential of the forward process is strongly convex. It further incorporates discretization to provide averaged first‑order stationarity bounds—sampling analogues of averaged gradient‑norm guarantees in nonconvex optimization—for samplers of both diffusion models. These results highlight a unique advantage of SDE‑based reverse diffusion over ODE‑based approaches, offering local convexity‑free certificates that ensure score consistency rather than global mode weights.

By Zhifeng Chen, Chenyang Jiang, Yazhen Wang