arXiv Machine Learning By Berk Tinaz, Changzhi Xie, Mahdi Soltanolkotabi

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks

Read the original on arXiv Machine Learning →

arXiv:2606. 04476v1 Announce Type: new Abstract: In this paper, we study the gradient descent dynamics for jointly training both layers of a one-hidden-layer ReLU network to fit a linear target function.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
4d ago

AYLA: Architecting a loss landscape in shallow neural networks to accelerate feature recovery

AYLA is a loss reparameterization framework that applies a sigmoid‑controlled power‑law transformation to the empirical loss, dynamically adjusting gradient magnitudes without changing stationary points or optimal solutions. By reshaping optimization trajectories, AYLA accelerates descent in flat or saddle‑dominated regions and stabilizes late‑stage training, leading to improved feature recovery in two‑layer tanh networks on synthetic Gaussian data. Experiments show enhanced weight alignment, neuron similarity, activation correlation, and richer internal representations, while mitigating rank collapse and promoting a transition from lazy to active feature‑learning regimes.

By Behnam Gheshlaghi, Shahin Atakishiyev