arXiv Machine Learning By Huijie Zhang, Jinfan Zhou, Yifu Lu, Minzhe Guo, Peng Wang, Liyue Shen, Qing Qu

The Emergence of Reproducibility and Generalizability in Diffusion Models

Read the original on arXiv Machine Learning →

arXiv:2310. 05264v5 Announce Type: replace Abstract: In this work, we investigate an intriguing and prevalent phenomenon of diffusion models which we term as "consistent model reproducibility": given the same starting noise input and a deterministic sampler, different diffusion models often yield remarkably similar outputs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 10

Noise in Diffusion Models Is a Learnable Input

The paper argues that the concrete random noise used in diffusion models is not merely a passive perturbation but a learnable input that can be exploited by the model. By analyzing how clean data and realized noise jointly form the noisy input, the authors show that the model can learn regularities in the data or in the noise structure, and that these two routes can interact. Experiments on MNIST and CIFAR‑10 using pseudorandom streams demonstrate that structured‑noise training can reduce prediction loss, but this advantage disappears when test noise is replaced with IID noise, indicating that the learned dependence is tied to the specific noise structure.

By Shengzhi Deng, Chenqi Ye, Yanze Guo
arXiv AI
3d ago

Distilling Diffusion Score Discrepancy for Efficient Training Data Attribution

The paper introduces a new method for training data attribution in diffusion models called TID, which uses a local score discrepancy measure and can be estimated without retraining. It further distills this approach into TIDE, a forward‑only student that reproduces the teacher’s rankings using internal activations, achieving comparable accuracy at dramatically lower query cost. Experiments on CIFAR‑10, ArtBench‑10, and MS‑COCO show that TID outperforms existing methods and TIDE attributes samples in milliseconds, faster than generation itself.

By Shixuan Liu, Joan Serr\`a, Kin Wai Cheuk, Jinju Kim, Woosung Choi, Yukara Ikemiya, Wei-Hsiang Liao, Jiaqi W. Ma, Yuki Mitsufuji
arXiv Machine Learning
Jun 9

Evaluating the Representation Space of Diffusion Models via Self-Supervised Principles

arXiv:2606. 09718v1 Announce Type: new Abstract: Diffusion models have demonstrated remarkable generative capabilities and have also emerged as powerful self-supervised representation learners, yet the connection between these two abilities remains less explored.

By Xiao Li, Yixuan Jia, Zekai Zhang, Xiang Li, Lianghe Shi, Jinxin Zhou, Zhihui Zhu, Liyue Shen, Qing Qu
arXiv AI
6d ago

Does Uniform Discrete Diffusion Need Time?

Uniform discrete diffusion models (UDMs) typically rely on explicit time conditioning, yet this study finds that such conditioning is often unnecessary in practice. While the population‑optimal UDM predictor generally depends on time—controlling how much the model should trust the observed context—the dependence becomes negligible in finite‑data language settings. Empirical results show that trained language UDMs exhibit limited time sensitivity across most of the diffusion trajectory, and time‑agnostic predictors can match or outperform time‑conditioned models on various datasets and training objectives.

By Chunsan Hong, Chieh-Hsin Lai, Satoshi Hayakawa, Yuhta Takida, Jong Chul Ye, Yuki Mitsufuji