arXiv Machine Learning By Ning Yang, Yikuan Zhang, Qi Ouyang, Chao Tang, Yuhai Tu

Noise-Driven Exploration and Transient Freezing Select Flat Minima in Stochastic Gradient Descent

Read the original on arXiv Machine Learning →

arXiv:2601. 10962v2 Announce Type: replace Abstract: Stochastic gradient descent (SGD) is central to deep learning, yet the dynamical origin of its preference for flatter, more generalizable solutions remains unclear.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
5d ago

AYLA: Architecting a loss landscape in shallow neural networks to accelerate feature recovery

AYLA is a loss reparameterization framework that applies a sigmoid‑controlled power‑law transformation to the empirical loss, dynamically adjusting gradient magnitudes without changing stationary points or optimal solutions. By reshaping optimization trajectories, AYLA accelerates descent in flat or saddle‑dominated regions and stabilizes late‑stage training, leading to improved feature recovery in two‑layer tanh networks on synthetic Gaussian data. Experiments show enhanced weight alignment, neuron similarity, activation correlation, and richer internal representations, while mitigating rank collapse and promoting a transition from lazy to active feature‑learning regimes.

By Behnam Gheshlaghi, Shahin Atakishiyev