arXiv AI By Tian Ding, Dawei Li, Ruoyu Sun

A Geometric Characterization of the Stationary Plateau for Two-Layer Neural Networks

Read the original on arXiv AI →

arXiv:2606. 04327v1 Announce Type: cross Abstract: We investigate the geometric structure of stationary plateaus that arise in the loss landscape of two-layer neural networks with smooth activation functions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 17

Beyond Quadratic Loss: The Stability Phase Diagram of Adam

The paper studies how Adam’s two momentum timescales, β1 and β3, influence loss spikes during neural‑network training. By mapping training dynamics across the (β1,β3) plane, the authors find an approximately linear boundary, 1-β3 = C(1-β1), that separates spiky from non‑spiky behavior, with the coefficient C linked to the effective loss exponent in superquadratic loss functions. They also show that confident cross‑entropy losses create a core–wall landscape that behaves superquadratically at the scale of an optimizer update, explaining the observed spikes.

By Gaoxiang Tang, Huanran Chen, Ziming Liu
arXiv Machine Learning
Jul 16

How the Hessian-Spectrum of Neural Networks Depends on Data

arXiv:2607. 13631v1 Announce Type: new Abstract: The Hessian matrix is an important quantity of interest when it comes to studying the loss landscape and optimization dynamics in deep learning, as well as designing measures of generalization, second-order learning algorithms, etc.

By Jasraj Singh, Enea Monzio Compagnoni, Antonio Orvieto