arXiv Machine Learning By Mo Zhou, Weihang Xu, Simon S. Du, Maryam Fazel

Learning Orthogonal Multi-Index Models Beyond Small Initialization: Incremental Learning, Competitive Dynamics and Symmetry

Read the original on arXiv Machine Learning →

The paper investigates how two‑layer polynomial‑width neural networks learn orthogonal multi‑index targets under standard initialization. It shows that incremental learning still occurs: the loss decreases sequentially following the Hermite expansion, with lower‑order components learned first. The dynamics also exhibit a competitive reallocation of parameter mass, shifting into the target subspace and concentrating on aligned neurons. The analysis uses a symmetry‑based finite‑width approximation and demonstrates that vanilla gradient descent displays the same qualitative behavior.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 4

High-Dimensional Learning Dynamics of Attention-Indexed Models

The paper investigates the training dynamics of attention mechanisms in high-dimensional settings, focusing on attention-indexed models that encompass multi-layer and multi-head architectures. It shows that while the loss landscape can be described by a finite set of trace order parameters, the online stochastic gradient descent dynamics involve an infinite hierarchy of matrix moments that can be accurately approximated by a finite truncated system. The study further reveals that the choice of attention parameterization acts as an implicit bias: untied attention can get trapped in uninformative states, whereas tied attention induces symmetry breaking and enables weak recovery with θ(d² log d) samples, and untied attention exhibits a fast-slow dynamic leading to weak recovery when symmetry is broken.

By Yizhou Xu, Margarita Sagitova, Lenka Zdeborov\'a, Florent Krzakala