The paper proves that for discrete diffusion models using uniform or remasking forward processes, an adaptive sampler based on a leave‑one‑out denoiser can achieve sampling error proportional to the score‑estimation error plus a small tolerance. The required number of discretization steps scales with the dual total correlation of the target distribution, not directly with the ambient dimension. This result shows that sampling complexity is governed by the intrinsic dependence structure of the distribution, and the authors provide an information‑theoretic analysis linking discretization error to mutual information between coordinates.
By Daniil Dmitriev, Zhihan Huang, Yuting Wei
arXiv:2607. 08757v1 Announce Type: cross Abstract: Score matching controls average error under the forward marginals, but a discretized reverse-time sampler evaluates the learned score along its own trajectory.
By Yiwei Zhou
arXiv:2506. 11378v3 Announce Type: replace Abstract: Sampling in score-based diffusion models can be performed by solving either a reverse-time stochastic differential equation (SDE) parameterized by an arbitrary stochasticity function or a probability flow ODE, corresponding to setting this stochasticity function to zero.
By Bernardo P. Schaeffer, Ricardo M. S. Rosa, Glauco Valle
arXiv:2512. 24152v2 Announce Type: replace-cross Abstract: Sampling based on score diffusions has led to striking empirical results, and has attracted considerable attention from various research communities.
By M. J. Wainwright
The paper introduces Hessian-free high-resolution (HFHR) dynamics, an extension of underdamped Langevin dynamics that incorporates reversible position diffusion for sampling in machine learning. It provides an explicit quantitative contraction rate under a position Poincaré inequality, weighted Hessian and Laplacian bounds, and a compact Sobolev embedding, even when the potential is non‑convex. For the HFHR Monte Carlo algorithm, a path‑space Girsanov argument yields a non‑asymptotic convergence bound and an explicit iteration complexity in total variation distance, improving on previous HFHR results and demonstrating benefits of a positive diffusion parameter through numerical experiments.
By Wujun Lv, Xiaoyu Wang, Yingli Wang, Lingjiong Zhu
arXiv:2501. 12982v3 Announce Type: replace-cross Abstract: This paper investigates how diffusion generative models leverage (unknown) low-dimensional structure to accelerate sampling.
By Jiadong Liang, Zhihan Huang, Yuxin Chen
arXiv:2607. 23226v1 Announce Type: new Abstract: Despite the empirical success of score-based diffusion models, a complete theoretical understanding of how finite-sample learning, network parameterization, and numerical discretization jointly dictate generative quality remains underdeveloped.
By Jinshu Huang, Yiming Jiang, Chunlin Wu
arXiv:2111. 10722v4 Announce Type: replace-cross Abstract: We propose a novel deterministic sampling method, EVI-MMD, to approximate a target distribution $\rho^*$ by minimizing the kernel discrepancy, also known as the Maximum Mean Discrepancy (MMD).
By Yindong Chen, Yiwei Wang, Lulu Kang, Chun Liu
arXiv:2502. 08834v4 Announce Type: replace-cross Abstract: Deep generative models based on neural differential equations have become state-of-the-art for many generation tasks.
By Zander W. Blasingame, Chen Liu
arXiv:2606. 06179v1 Announce Type: cross Abstract: Score-based diffusion models are typically trained by minimizing the $L^2$ score matching error, and standard theoretical analyses rely on this quantity to bound the sampling discrepancy between the learned and target distributions.
By Na\"il B. Khelifa, Richard E. Turner, Ramji Venkataramanan
arXiv:2602. 15008v2 Announce Type: replace Abstract: Diffusion models over discrete spaces have recently shown striking empirical success, yet their theoretical foundations remain incomplete.
By Daniil Dmitriev, Zhihan Huang, Yuting Wei
The paper presents a non‑asymptotic analysis of Markov chain Monte Carlo (MCMC) algorithms that learn and apply a preconditioner based on either the target covariance or the expected Hessian of the target potential. It compares the finite‑time computational costs of these preconditioned schemes with unpreconditioned counterparts, providing guarantees for algorithms such as the Unadjusted Langevin Algorithm (ULA) and the proximal sampler. The analysis relies on a contraction assumption in the Wasserstein‑2 distance to formalize approximate independence and bridge modern MCMC theory with classical effective sample size heuristics.
By Max Hird, Florian Maire, Jeffrey Negrea