arXiv:2601. 22450v2 Announce Type: replace-cross Abstract: Masked Diffusion Language Models have recently emerged as a powerful generative paradigm, yet their generalization properties remain understudied compared to their auto-regressive counterparts.
By Jianhao Huang, Baharan Mirzasoleiman
arXiv:2606. 13796v1 Announce Type: cross Abstract: Recursive training of generative models on their own outputs can lead to model collapse, a compounding drift away from the true data distribution.
By Na\"il B. Khelifa, Richard E. Turner, Ramji Venkataramanan
arXiv:2310. 05264v5 Announce Type: replace Abstract: In this work, we investigate an intriguing and prevalent phenomenon of diffusion models which we term as "consistent model reproducibility": given the same starting noise input and a deterministic sampler, different diffusion models often yield remarkably similar outputs.
By Huijie Zhang, Jinfan Zhou, Yifu Lu, Minzhe Guo, Peng Wang, Liyue Shen, Qing Qu
arXiv:2602. 02908v2 Announce Type: replace-cross Abstract: Diffusion models trained on different, non-overlapping subsets of a dataset often produce strikingly similar outputs when given the same noise seed.
By Binxu Wang, Jacob Zavatone-Veth, Cengiz Pehlevan
arXiv:2607. 07967v1 Announce Type: cross Abstract: Diffusion-based policies have recently emerged as powerful policy parameterizations for reinforcement learning, representing state-conditioned action distributions as terminal laws of diffusion processes with parameterized drifts.
By Viet Vu, Renyuan Xu, Jiacheng Zhang, Yufei Zhang
arXiv:2608. 02575v1 Announce Type: new Abstract: Diffusion models rely on stochastic inputs, yet on finite-precision hardware, the "randomness" they consume is realized as deterministic numerical orbits generated by pseudorandom rules.
By Shengzhi Deng, Chenqi Ye, Yanze Guo
arXiv:2608. 01833v1 Announce Type: cross Abstract: Grokking is a striking phenomenon in neural network training, where a model can undergo a prolonged period of pure memorization before abrupt generalization.
By Lai Shun Chan, Xiaotian Zhang, Yue Shang, Ge Zhang, Entao Yang
arXiv:2607. 02137v1 Announce Type: cross Abstract: We study timestep allocation for score-based diffusion sampling, where a learned reverse-time dynamics is discretized on a finite grid.
By Yilie Huang, Wenpin Tang, Xun Yu Zhou
arXiv:2602. 05533v3 Announce Type: replace Abstract: We study conditional generation in diffusion models under hard constraints, where generated samples must satisfy prescribed events with probability one.
By Zhengyi Guo, Wenpin Tang, Renyuan Xu
arXiv:2606. 17192v1 Announce Type: new Abstract: This paper develops constrained diffusion models with primal-dual inference (PDI) to sample from optimal distributions of entropy-regularized optimization problems with \emph{average} constraints.
By Samar Hadou, Yigit Berkay Uslu, Alejandro Ribeiro
arXiv:2602. 02241v2 Announce Type: replace Abstract: Entropic optimal transport (EOT) in continuous spaces with quadratic cost is a classical tool for solving the domain translation problem.
By Roman Dyachenko, Nikita Gushchin, Kirill Sokolov, Petr Mokrov, Evgeny Burnaev, Alexander Korotin
arXiv:2604. 26985v2 Announce Type: replace-cross Abstract: Masked diffusion models (MDMs) generate discrete sequences by iterative denoising under an absorbing masking process.
By Michael Cardei, Huu Binh Ta, Ferdinando Fioretto