arXiv Machine Learning By Simon Kuang, Kyle Chickering, Xinfan Lin

Stable initialization without the CLT

Read the original on arXiv Machine Learning →

The paper introduces a new method called uniform‑phase initialization for deep neural networks with sine activations, eliminating the need for the Central Limit Theorem and fully decoupling layers. This approach avoids distributional approximation errors and coupling between layers, leading to stable weight initialization. Experiments show that models using this initialization outperform state‑of‑the‑art methods on image and audio fitting tasks and remain competitive without tuning, while also supporting μP width scaling.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 16

How Controlling the Variance can Improve Training Stability of Sparsely Activated DNNs and CNNs

arXiv:2602. 05779v2 Announce Type: replace Abstract: The Edge-of-Chaos (EoC) theory developed for the random initialization of deep networks allows more efficient training by both preserving information in the initial outputs of the network and minimising exploding or vanishing gradients through characterisation of the intermediate layers as Gaussian processes.

By Emily Dent, Jared Tanner
arXiv Machine Learning
Sep 11

Learning Orthogonal Multi-Index Models Beyond Small Initialization: Incremental Learning, Competitive Dynamics and Symmetry

The paper investigates how two‑layer polynomial‑width neural networks learn orthogonal multi‑index targets under standard initialization. It shows that incremental learning still occurs: the loss decreases sequentially following the Hermite expansion, with lower‑order components learned first. The dynamics also exhibit a competitive reallocation of parameter mass, shifting into the target subspace and concentrating on aligned neurons. The analysis uses a symmetry‑based finite‑width approximation and demonstrates that vanilla gradient descent displays the same qualitative behavior.

By Mo Zhou, Weihang Xu, Simon S. Du, Maryam Fazel