The paper models the dynamics of Stochastic Gradient Descent (SGD) as a percolation process, showing that architectural symmetries cause subnetworks to merge in discrete blocks rather than sequentially. These structural transitions produce variance spikes in a macroscopic order parameter, analogous to physical phase transitions. The authors also demonstrate that this trapping mechanism and its scaling cascade apply to Adam and AdamW under a heavy‑tailed noise model.
By Sai Niranjan Ramachandran, Suvrit Sra
arXiv:2607. 16761v1 Announce Type: cross Abstract: Dropout and Random Gradient Masking (RaM) are two training techniques used to improve performance in deep learning.
By Javier Maass, L\'ena\"ic Chizat
The paper presents a supervised, scale‑shared neural architecture that learns a coarse‑graining rule for two‑dimensional site percolation. By recursively applying this rule, the model generates a latent field from which the crossing probability is predicted and a fine‑graining decoder reconstructs the largest‑cluster mask. Trained only on small lattices, the network extrapolates to larger systems, accurately recovers the spanning cluster, and reproduces finite‑size scaling near the critical point, demonstrating that the latent representation captures critical fluctuations and scale‑dependent flows consistent with renormalization‑group theory.
By Anaclara Alvez, Luca Camagna, Sergio Chibbaro, Cyril Furtlehner, Fran\c{c}ois Landes, Gianluca Manzan, Lorenzo Mensi
arXiv:2606. 31282v1 Announce Type: new Abstract: Modern deep neural networks often contain far more parameters than needed to fit their training data, yet they achieve impressive generalization.
By Ari Pakman, Lior Kreimer, Yakir Berchenko
The paper introduces a supervised, scale‑shared neural architecture for two‑dimensional site percolation, implementing a neural renormalization group flow. The model recursively applies a learned coarse‑graining rule across scales, producing a latent field that predicts crossing probability and a fine‑graining decoder that reconstructs the largest‑cluster mask. Trained only on small lattices, it extrapolates to larger systems, accurately recovers the spanning cluster, and yields observables that follow expected finite‑size scaling near the critical point, highlighting the importance of critical fluctuations in the latent representation.
arXiv:2507. 14159v2 Announce Type: replace-cross Abstract: Predicting critical phenomena from limited labeled data remains a challenging task in statistical physics.
By Shanshan Wang, Dian Xu, Jianmin Shen, Feng Gao, Wei Li, Weibing Deng