arXiv Machine Learning

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks

arXiv:2606. 04476v1 Announce Type: new Abstract: In this paper, we study the gradient descent dynamics for jointly training both layers of a one-hidden-layer ReLU network to fit a linear target function.