arXiv Machine Learning

Pseudorandom Streams within Diffusion Models Act as Learnable Inputs That Affect Generation Quality

arXiv:2608. 02575v1 Announce Type: new Abstract: Diffusion models rely on stochastic inputs, yet on finite-precision hardware, the "randomness" they consume is realized as deterministic numerical orbits generated by pseudorandom rules.

arXiv Machine Learning
Jun 10

The Emergence of Reproducibility and Generalizability in Diffusion Models

arXiv:2310. 05264v5 Announce Type: replace Abstract: In this work, we investigate an intriguing and prevalent phenomenon of diffusion models which we term as "consistent model reproducibility": given the same starting noise input and a deterministic sampler, different diffusion models often yield remarkably similar outputs.

By Huijie Zhang, Jinfan Zhou, Yifu Lu, Minzhe Guo, Peng Wang, Liyue Shen, Qing Qu
arXiv Machine Learning
Sep 10

Noise in Diffusion Models Is a Learnable Input

The paper argues that the concrete random noise used in diffusion models is not merely a passive perturbation but a learnable input that can be exploited by the model. By analyzing how clean data and realized noise jointly form the noisy input, the authors show that the model can learn regularities in the data or in the noise structure, and that these two routes can interact. Experiments on MNIST and CIFAR‑10 using pseudorandom streams demonstrate that structured‑noise training can reduce prediction loss, but this advantage disappears when test noise is replaced with IID noise, indicating that the learned dependence is tied to the specific noise structure.

By Shengzhi Deng, Chenqi Ye, Yanze Guo
arXiv AI
3d ago

Distilling Diffusion Score Discrepancy for Efficient Training Data Attribution

The paper introduces a new method for training data attribution in diffusion models called TID, which uses a local score discrepancy measure and can be estimated without retraining. It further distills this approach into TIDE, a forward‑only student that reproduces the teacher’s rankings using internal activations, achieving comparable accuracy at dramatically lower query cost. Experiments on CIFAR‑10, ArtBench‑10, and MS‑COCO show that TID outperforms existing methods and TIDE attributes samples in milliseconds, faster than generation itself.

By Shixuan Liu, Joan Serr\`a, Kin Wai Cheuk, Jinju Kim, Woosung Choi, Yukara Ikemiya, Wei-Hsiang Liao, Jiaqi W. Ma, Yuki Mitsufuji
arXiv Machine Learning
Aug 31

Improved off-policy training of diffusion samplers

The paper investigates training diffusion models to sample from distributions defined by unnormalized densities or energy functions. It benchmarks various diffusion-structured inference techniques, including simulation-based variational methods and off-policy approaches such as continuous generative flow networks, highlighting their relative strengths and challenging some prior claims. Additionally, the authors introduce a new exploration strategy for off-policy methods that employs local search in the target space with a replay buffer, demonstrating improved sample quality across multiple target distributions.

By Marcin Sendera, Minsu Kim, Sarthak Mittal, Pablo Lemos, Luca Scimeca, Jarrid Rector-Brooks, Alexandre Adam, Yoshua Bengio, Esmeralda S. Whitammer
arXiv AI
Sep 3

Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data

The paper investigates when language diffusion models, specifically Uniform-based Discrete Diffusion Models (UDDMs), shift from memorizing training data to generalizing to new data. It shows that UDDMs act as associative memories, forming basins of attraction around stored examples without requiring an explicit energy function. By measuring token recovery and conditional entropy, the authors identify a sharp transition governed by training set size, where memorization (vanishing entropy) gives way to generalization (finite entropy).

By Bao Pham, Mohammed J. Zaki, Luca Ambrogioni, Dmitry Krotov, Matteo Negri
arXiv Machine Learning
Sep 22

Optimizers for Diffusion Models: A Controlled Benchmark

The paper introduces the first controlled benchmark for optimizers in discrete diffusion models, evaluating seven optimizers (AdamW, Lion, Muon, SOAP, MARS, MARS‑M, Schedule‑Free) across four diffusion formulations: masked diffusion on text8, uniform diffusion on QM9 and LM1B, and Gaussian diffusion on CelebA‑64. Each optimizer undergoes the same search protocol and is retrained with full budget and multiple seeds, revealing that AdamW, while strong, is not universally optimal and that optimizers validated on autoregressive language models (Muon, MARS‑M, SOAP) can outperform tuned AdamW on certain tasks.

By Arman Bolatov, Egor Shulgin, David Li, Abduragim Shtanchaev, Sebastian U. Stich, Maxim Panov, Eric Moulines, Peter Richt\'arik, Martin Tak\'a\v{c}