We introduce Sticky Jump Diffusions (SJDs), continuous-time Markov processes on $\mathbb R^d$ whose discrete anchors are token embeddings. In forward time, anchors release their mass at a hazard rate and the released mass diffuses in the continuous ambient space; time reversal couples a score-driven SDE with a sticky jump kernel whose rate and destination are fixed by flux balance with the forward law.
The article "Foundations of Diffusion Models in General State Spaces: A Self-Contained Introduction" presents a unified primer on diffusion models that applies to both continuous Euclidean data and discrete categorical structures. It develops discrete-time forward noising via Markov kernels and learned reverse dynamics, and connects these to continuous-time limits such as stochastic differential equations in ρ^d and continuous-time Markov chains on finite alphabets, deriving the corresponding Fokker–Planck and master equations. The work also shows how different forward corruption choices—Gaussian processes for continuous spaces and structured categorical transition kernels for discrete spaces—affect reverse dynamics and the evidence lower bound used in training, offering a layered exposition for newcomers, practitioners, and experts alike.
By Vincent Pauline, Tobias H\"oppe, Kirill Neklyudov, Alexander Tong, Stefan Bauer, Andrea Dittadi
arXiv:2607. 05381v1 Announce Type: cross Abstract: What does a discrete diffusion model learn: a denoiser, a score ratio, or a bridge plug-in predictor?
By Rodrigo Casado Noguerales, Bernhard Sch\"olkopf, Thomas Hofmann, Aran Raoufi
arXiv:2604.27443v3 Announce Type: replace
Abstract: Generating continuous-time, continuous-space stochastic processes (e.g., videos, weather forecasts) conditioned on partial observations (e.g., firs...
By Gabe Guo, Thanawat Sornwanee, Lutong Hao, Elon Litman, Stefano Ermon, Jose Blanchet
arXiv:2604. 26985v2 Announce Type: replace-cross Abstract: Masked diffusion models (MDMs) generate discrete sequences by iterative denoising under an absorbing masking process.
By Michael Cardei, Huu Binh Ta, Ferdinando Fioretto
arXiv:2606. 02232v1 Announce Type: new Abstract: Learning a Markov transition model is not merely conditional density estimation: the learned object must be a valid transition kernel before it is iterated in downstream dynamics.
By Ao Xu
arXiv:2607. 24507v1 Announce Type: cross Abstract: Existing methods mainly adapt pretrained autoregressive (AR) language models to masked diffusion, whereas we directly adapt them to uniform-noise diffusion, where every token remains editable during sampling.
By Xiaoyi Jiang, Jingyuan Li, Yixuan Jiang, Wei Liu, Yi Zhu, Zuoqiang Shi, Pipi Hu
arXiv:2606. 24000v1 Announce Type: new Abstract: We introduce cyclic denoising -- repeated forward and reverse diffusion at controlled noise amplitudes -- as an extraction attack for image diffusion models.
By Rishabh Sharma, Stefano Martiniani
arXiv:2608.23916v1 Announce Type: new
Abstract: Denoising score matching trains diffusion models by regressing onto a conditional score, although generation ultimately requires the marginal score. Th...
By Avinash Raju, Kai Zhang
arXiv:2512. 02657v2 Announce Type: replace-cross Abstract: Real-world deployment of text-to-image diffusion models requires continual concept removal as new privacy, copyright, or safety obligations arise over time.
By Naveen George, Naoki Murata, Yuhta Takida, Konda Reddy Mopuri, Yuki Mitsufuji
arXiv:2603. 15384v2 Announce Type: replace-cross Abstract: We improve and extend persistence spheres, introduced in~\cite{pegoraro2025persistence}.
By Matteo Pegoraro
arXiv:2608.10615v2 Announce Type: replace
Abstract: Discrete diffusion models for categorical generation are defined by a corruption kernel, which determines the intermediate state space and the asso...
By Jinya Sakurai, Patrick Pynadath, Satoshi Hayakawa, Jaehong Yoon, Xulei Yang, Nancy F. Chen, Xun Xu