The paper argues that the concrete random noise used in diffusion models is not merely a passive perturbation but a learnable input that can be exploited by the model. By analyzing how clean data and realized noise jointly form the noisy input, the authors show that the model can learn regularities in the data or in the noise structure, and that these two routes can interact. Experiments on MNIST and CIFAR‑10 using pseudorandom streams demonstrate that structured‑noise training can reduce prediction loss, but this advantage disappears when test noise is replaced with IID noise, indicating that the learned dependence is tied to the specific noise structure.
By Shengzhi Deng, Chenqi Ye, Yanze Guo
arXiv:2606. 13801v1 Announce Type: new Abstract: Neural responses in cortex exhibit substantial trial-to-trial variability in response to repeated stimuli, while peripheral sensory neurons respond far more consistently, leading many to wonder whether stochasticity may carry meaning.
By Robin Preble, Praveen Venkatesh, Stefan Mihalas, Kameron Decker Harris
arXiv:2608. 05464v1 Announce Type: cross Abstract: The pruning of network connections is key to brain function but, despite its importance, there exist few biologically-plausible pruning rules with demonstrated good performance.
By Sanjith Senthil, Rishidev Chaudhuri
The paper examines how varying input noise characteristics—type, scale, and complexity—affect neural network robustness in geophysical tasks such as first break picking and denoising. By training models on fixed noise settings and testing them on both seen and unseen noise scenarios, the study constructs a robustness matrix that reveals how larger noise scales improve generalization and how aligning noise type with task complexity and architecture maximizes performance. Training with compound noise mixtures further mitigates weaknesses of single-noise training, acting as an implicit regularizer that enhances robustness under out‑of‑distribution conditions.
By Salma Alsinan, Maksim Makarenko, Sixiu Liu, Ali Aldawood, Ibrahim Hoteit
arXiv:2602. 14885v2 Announce Type: replace-cross Abstract: Recurrent neural networks (RNNs) provide a theoretical framework for understanding computation in biological neural circuits, yet classical results, such as Hopfield's model of associative memory, rely on symmetric connectivity that restricts network dynamics to gradient-like flows.
By Ram\'on Nartallo-Kaluarachchi, Renaud Lambiotte, Alain Goriely
arXiv:2605.08144v2 Announce Type: replace-cross
Abstract: Training a diffusion model involves two sources of randomness for each data sample: the timestep and the Gaussian noise realization. The time...
By Haokai Zhao, Da Xing, Hanqun Cao, Tinson Xu, Xinyu Xiang, Yanchao Li, Xiangru Tang, Hongbin Lin, Zehong Wang, Kuan Pang, Peng Xia, Molei Tao, Li Erran Li, Aditya Joshi, Jure Leskovec, Fang Wu
The pruning of network connections is key to brain function but, despite its importance, there exist few biologically-plausible pruning rules with demonstrated good performance. In this work we evaluate noise-prune, a recently introduced unsupervised local pruning rule for recurrent networks that uses noisy fluctuations to determine the importance of connections.
arXiv:2607. 14466v1 Announce Type: new Abstract: Noise injection is a well-known technique in stochastic optimization.
By Matt L. Wiemann, Peter Melchior, Andrew K. Saydjari
arXiv:2602. 04078v2 Announce Type: replace-cross Abstract: Deep learning has achieved remarkable success across a wide range of domains, significantly expanding the frontiers of what is achievable in artificial intelligence.
By R\'ois\'in Luo
arXiv:2607. 28185v1 Announce Type: new Abstract: Oversmoothing is a fundamental limitation of deep graph neural networks (GNNs), where repeated message passing causes node representations to become increasingly similar, eventually collapsing toward a low-dimensional subspace.
By Mostafa Haghir Chehreghani
arXiv:2607. 12360v1 Announce Type: new Abstract: The cooldown phase of a warmup-stable-decay (WSD) learning-rate schedule, now a default in large-model pretraining, lowers the final training loss in some settings and does nothing in others.
By Subham Singh, Ashutosh Mishra, Subha Raut
arXiv:2607. 16761v1 Announce Type: cross Abstract: Dropout and Random Gradient Masking (RaM) are two training techniques used to improve performance in deep learning.
By Javier Maass, L\'ena\"ic Chizat