arXiv AI

NoiseRater: Meta-Learned Noise Valuation for Diffusion Model Training

arXiv Machine Learning
Sep 10

Noise in Diffusion Models Is a Learnable Input

The paper argues that the concrete random noise used in diffusion models is not merely a passive perturbation but a learnable input that can be exploited by the model. By analyzing how clean data and realized noise jointly form the noisy input, the authors show that the model can learn regularities in the data or in the noise structure, and that these two routes can interact. Experiments on MNIST and CIFAR‑10 using pseudorandom streams demonstrate that structured‑noise training can reduce prediction loss, but this advantage disappears when test noise is replaced with IID noise, indicating that the learned dependence is tied to the specific noise structure.

By Shengzhi Deng, Chenqi Ye, Yanze Guo
arXiv Machine Learning
Jun 10

The Emergence of Reproducibility and Generalizability in Diffusion Models

arXiv:2310. 05264v5 Announce Type: replace Abstract: In this work, we investigate an intriguing and prevalent phenomenon of diffusion models which we term as "consistent model reproducibility": given the same starting noise input and a deterministic sampler, different diffusion models often yield remarkably similar outputs.

By Huijie Zhang, Jinfan Zhou, Yifu Lu, Minzhe Guo, Peng Wang, Liyue Shen, Qing Qu
arXiv AI
Jun 2

Paradoxical noise preference in RNNs

arXiv:2601. 04539v2 Announce Type: replace-cross Abstract: In recurrent neural networks (RNNs) used to model biological neural networks, noise is typically introduced during training to emulate biological variability and regularize learning.

By Noah Eckstein, Manoj Srinivasan
arXiv AI
6d ago

Does Uniform Discrete Diffusion Need Time?

Uniform discrete diffusion models (UDMs) typically rely on explicit time conditioning, yet this study finds that such conditioning is often unnecessary in practice. While the population‑optimal UDM predictor generally depends on time—controlling how much the model should trust the observed context—the dependence becomes negligible in finite‑data language settings. Empirical results show that trained language UDMs exhibit limited time sensitivity across most of the diffusion trajectory, and time‑agnostic predictors can match or outperform time‑conditioned models on various datasets and training objectives.

By Chunsan Hong, Chieh-Hsin Lai, Satoshi Hayakawa, Yuhta Takida, Jong Chul Ye, Yuki Mitsufuji
arXiv Machine Learning
Sep 22

Leveraging Inference-Time Compute for Diffusion Models via Global Scheduling of Denoising Trajectories

The paper studies how to allocate a fixed computational budget across the denoising steps of diffusion models to improve sample quality at deployment. It shows that the expected benefit of evaluating multiple candidates at a step can be decomposed into a step‑specific sensitivity and a universal sample‑size factor, and that the optimal allocation follows a water‑filling structure. Experiments demonstrate that this allocation achieves the same quality as a uniform strategy while reducing function evaluations by 20–50%.

By Yuan Cao, Yifu Tang, Hangqi Li, Zeyu Zheng
arXiv Machine Learning
Sep 22

Optimizers for Diffusion Models: A Controlled Benchmark

The paper introduces the first controlled benchmark for optimizers in discrete diffusion models, evaluating seven optimizers (AdamW, Lion, Muon, SOAP, MARS, MARS‑M, Schedule‑Free) across four diffusion formulations: masked diffusion on text8, uniform diffusion on QM9 and LM1B, and Gaussian diffusion on CelebA‑64. Each optimizer undergoes the same search protocol and is retrained with full budget and multiple seeds, revealing that AdamW, while strong, is not universally optimal and that optimizers validated on autoregressive language models (Muon, MARS‑M, SOAP) can outperform tuned AdamW on certain tasks.

By Arman Bolatov, Egor Shulgin, David Li, Abduragim Shtanchaev, Sebastian U. Stich, Maxim Panov, Eric Moulines, Peter Richt\'arik, Martin Tak\'a\v{c}