The paper models the dynamics of Stochastic Gradient Descent (SGD) as a percolation process, showing that architectural symmetries cause subnetworks to merge in discrete blocks rather than sequentially. These structural transitions produce variance spikes in a macroscopic order parameter, analogous to physical phase transitions. The authors also demonstrate that this trapping mechanism and its scaling cascade apply to Adam and AdamW under a heavy‑tailed noise model.
By Sai Niranjan Ramachandran, Suvrit Sra
arXiv:2607. 16761v1 Announce Type: cross Abstract: Dropout and Random Gradient Masking (RaM) are two training techniques used to improve performance in deep learning.
By Javier Maass, L\'ena\"ic Chizat
The paper presents a supervised, scale‑shared neural architecture that learns a coarse‑graining rule for two‑dimensional site percolation. By recursively applying this rule, the model generates a latent field from which the crossing probability is predicted and a fine‑graining decoder reconstructs the largest‑cluster mask. Trained only on small lattices, the network extrapolates to larger systems, accurately recovers the spanning cluster, and reproduces finite‑size scaling near the critical point, demonstrating that the latent representation captures critical fluctuations and scale‑dependent flows consistent with renormalization‑group theory.
By Anaclara Alvez, Luca Camagna, Sergio Chibbaro, Cyril Furtlehner, Fran\c{c}ois Landes, Gianluca Manzan, Lorenzo Mensi
arXiv:2606. 31282v1 Announce Type: new Abstract: Modern deep neural networks often contain far more parameters than needed to fit their training data, yet they achieve impressive generalization.
By Ari Pakman, Lior Kreimer, Yakir Berchenko
The paper introduces a supervised, scale‑shared neural architecture for two‑dimensional site percolation, implementing a neural renormalization group flow. The model recursively applies a learned coarse‑graining rule across scales, producing a latent field that predicts crossing probability and a fine‑graining decoder that reconstructs the largest‑cluster mask. Trained only on small lattices, it extrapolates to larger systems, accurately recovers the spanning cluster, and yields observables that follow expected finite‑size scaling near the critical point, highlighting the importance of critical fluctuations in the latent representation.
arXiv:2507. 14159v2 Announce Type: replace-cross Abstract: Predicting critical phenomena from limited labeled data remains a challenging task in statistical physics.
By Shanshan Wang, Dian Xu, Jianmin Shen, Feng Gao, Wei Li, Weibing Deng
arXiv:2510. 24616v4 Announce Type: replace-cross Abstract: For four decades statistical physics has been providing a framework to analyse neural networks.
By Jean Barbier, Francesco Camilli, Minh-Toan Nguyen, Mauro Pastore, Rudy Skerk
The paper investigates how stochastic gradient descent (SGD) selects specific functional decompositions when training a deep linear residual network to learn the identity function. Although many weight configurations minimize the population loss, SGD consistently prefers particular solutions, especially under anisotropic label noise or different parametrizations. The authors explain this bias using an entropic loss term that penalizes the expected squared norm of the minibatch gradient, analytically characterizing its minimizers and showing that trained networks align with these predictions.
By Andy Arditi, Weian Xie, David Bau, Liu Ziyin
The paper investigates how to reduce computation in neural networks by combining one‑shot magnitude pruning in a static setting with early exit in an adaptive setting. In a simplified single‑neuron model it proves a concentration theorem for pruning and introduces a conditional perceptron whose excess error decreases as a power of the compute gap, with the exponent increasing as partial and full computations align. The authors extend these results to deep networks, showing how pruning distortions accumulate with depth and deriving a compute‑accuracy trade‑off for frozen‑backbone early exit under a Gaussian process framework, with numerical simulations supporting the theoretical scaling laws.
By Erdem Koyuncu
arXiv:2608. 01833v1 Announce Type: cross Abstract: Grokking is a striking phenomenon in neural network training, where a model can undergo a prolonged period of pure memorization before abrupt generalization.
By Lai Shun Chan, Xiaotian Zhang, Yue Shang, Ge Zhang, Entao Yang
arXiv:2607. 16720v1 Announce Type: new Abstract: Understanding deep neural networks remains a central challenge in machine learning.
By Haruka Eshima, Makoto Yamada
arXiv:2602. 05779v2 Announce Type: replace Abstract: The Edge-of-Chaos (EoC) theory developed for the random initialization of deep networks allows more efficient training by both preserving information in the initial outputs of the network and minimising exploding or vanishing gradients through characterisation of the intermediate layers as Gaussian processes.
By Emily Dent, Jared Tanner