arXiv:2602. 10949v2 Announce Type: replace-cross Abstract: Effective initialization in deep networks requires an understanding of random neural networks.
By Constantin Kogler, Tassilo Schwarz, Samuel Kittle
arXiv:2506. 08764v3 Announce Type: replace Abstract: Deep neural networks are known to suffer from exploding or vanishing gradients as depth increases, a phenomenon closely tied to the spectral behavior of the input-output Jacobian.
By Benjamin Dadoun, Soufiane Hayou, Hanan Salam, Mohamed El Amine Seddik, Pierre Youssef
arXiv:2602. 05779v2 Announce Type: replace Abstract: The Edge-of-Chaos (EoC) theory developed for the random initialization of deep networks allows more efficient training by both preserving information in the initial outputs of the network and minimising exploding or vanishing gradients through characterisation of the intermediate layers as Gaussian processes.
By Emily Dent, Jared Tanner
arXiv:2606.06671v2 Announce Type: replace
Abstract: Existing implicit neural representation (INR) approaches suffer from stochastic initialization that does not guarantee consistent or high-quality p...
By Mohammed Alsakabi, Kejia Hu, John M. Dolan, Ozan K. Tonguz
arXiv:2502. 18959v4 Announce Type: replace Abstract: The architecture of a neural network and the choice of its activation function are both fundamental to its performance.
By Shijun Zhang, Hongkai Zhao, Yimin Zhong, Haomin Zhou
The paper investigates how two‑layer polynomial‑width neural networks learn orthogonal multi‑index targets under standard initialization. It shows that incremental learning still occurs: the loss decreases sequentially following the Hermite expansion, with lower‑order components learned first. The dynamics also exhibit a competitive reallocation of parameter mass, shifting into the target subspace and concentrating on aligned neurons. The analysis uses a symmetry‑based finite‑width approximation and demonstrates that vanilla gradient descent displays the same qualitative behavior.
By Mo Zhou, Weihang Xu, Simon S. Du, Maryam Fazel
arXiv:2608. 11970v1 Announce Type: new Abstract: The parity problem--deciding whether the number of ones in a binary vector is odd or even--remains challenging for standard neural networks due to linear inseparability and the need for global interactions.
By Daehwa Ko, Jaehyeon Kim, Seunghyun Ham, Jay Hoon Jung
arXiv:2606. 04754v1 Announce Type: new Abstract: Many striking phenomena in deep learning, such as linear mode connectivity and the structured behavior of training dynamics, are closely tied to parameter symmetries: transformations that leave the realized function unchanged.
By Vincent B\"urgin, Daniel Herbst, Ya-Wei Eileen Lin, Stefanie Jegelka
arXiv:2606. 04476v1 Announce Type: new Abstract: In this paper, we study the gradient descent dynamics for jointly training both layers of a one-hidden-layer ReLU network to fit a linear target function.
By Berk Tinaz, Changzhi Xie, Mahdi Soltanolkotabi
arXiv:2602. 19799v2 Announce Type: replace-cross Abstract: Despite recent algorithmic advances, we still lack principled ways to leverage the well-documented rescaling symmetries in ReLU neural network parameters.
By Arthur Lebeurrier, Titouan Vayer, R\'emi Gribonval
arXiv:2608. 04442v1 Announce Type: new Abstract: Robustness to natural corruptions remains a fundamental challenge for deep neural networks.
By Jiangang Yang, Wenhui Shi, Lu Hu, Jing Xing, Jian Liu
arXiv:2606. 02993v1 Announce Type: new Abstract: Understanding how structured internal structure emerges during neural network training is central to the study of deep learning.
By Jianliang He, Leda Wang, Fengzhuo Zhang, Siyu Chen, Zhuoran Yang