arXiv Machine Learning By Parsa Rangriz

Limit Theorems for Stochastic Gradient Descent in High-Dimensional Single-Layer Networks

Read the original on arXiv Machine Learning →

arXiv:2511. 02258v3 Announce Type: replace-cross Abstract: This paper studies the high-dimensional scaling limits of online stochastic gradient descent (SGD).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 3

Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance

The paper models the dynamics of Stochastic Gradient Descent (SGD) as a percolation process, showing that architectural symmetries cause subnetworks to merge in discrete blocks rather than sequentially. These structural transitions produce variance spikes in a macroscopic order parameter, analogous to physical phase transitions. The authors also demonstrate that this trapping mechanism and its scaling cascade apply to Adam and AdamW under a heavy‑tailed noise model.

By Sai Niranjan Ramachandran, Suvrit Sra
arXiv AI
3d ago

Mean--Fluctuation Dynamics at the Edge of Stability

The paper investigates gradient descent dynamics in the Edge of Stability regime, where a large learning rate causes persistent oscillations linked to improved generalization. It introduces a tractable continuous‑time mean–fluctuation model that couples the window‑averaged trajectory with its fluctuation covariance, derives this model rigorously from a sharp‑valley framework, and analyzes its stationary states and linear stability. The authors also extend the model to wide two‑layer networks, deriving a Wasserstein‑2 gradient flow for weights and fluctuations, proving well‑posedness, a mean‑field limit, and conditional convergence results, with numerical experiments illustrating the predictions and finite‑time limitations.

By Antonin Chodron de Courcel