arXiv Machine Learning

A Full Adam Theorem for Spectral Heavy-Tail Onset

arXiv:2609. 12996v1 Announce Type: new Abstract: We prove a full Adam theorem for spectral heavy-tail onset in a closed Gaussian Stein-Hermite teacher-student state-evolution model.

arXiv Machine Learning
Sep 14

Adapt or Forget: Provable Tradeoffs Between Adam and SGD in Nonstationary Optimization

The paper presents a theoretical study of Adam in non‑stationary stochastic optimization, distinguishing two regimes: Euclidean tracking under adaptive strong monotonicity and high‑probability projected stationarity for general smooth objectives. It derives finite‑time bounds that decompose into initialization, objective drift, first‑moment tracking error (β₁), and preconditioner perturbation (β₂), and characterizes burn‑in times for constant and step‑decay schedules. The analysis reveals a noise–drift tradeoff, showing that in noise‑dominated settings Adam’s adaptive mechanisms can improve guarantees, while in drift‑dominated settings they may worsen tracking, potentially making vanilla SGD preferable.

By Sharan Sahu, Abir Sarkar, Cameron J. Hogan, Martin T. Wells
arXiv Machine Learning
Sep 24

Linear RNN Scaling Laws: When Longer Sequences Beat More Sequences

The paper presents empirical scaling laws for autoregressive language models, linking prediction loss to model size, data size, and compute, and investigates their theoretical basis using a teacher–student linear RNN framework. In this tractable setting, a stable latent linear RNN generates trajectories while a sketched linear recurrent student is trained via full‑batch WSD gradient descent on next‑token prediction. The study derives explicit approximation, optimization, and statistical scaling laws that depend on the sketch dimension, number of trajectories, and trajectory length, revealing how different power‑law exponents for innovation and initialization covariances affect the rates and crossovers between regimes.

By Ziyan Chen, Zhongzhu Zhou, Peilin Liu, Ding-Xuan Zhou
arXiv Machine Learning
Sep 24

The Drift Contract: Spectral Updates for Depth-Robust Local Learning

The paper introduces the Drift Contract, a spectral update geometry for local learning that improves depth robustness and hyperparameter stability. By applying momentum orthogonalization with spectral step scaling to per‑layer updates, the authors achieve consistent performance across a wide range of widths and depths on CIFAR‑10 MLPs, outperforming local Adam and providing a per‑layer, input‑conditioned drift bound. The study also shows that the spectral geometry itself, rather than step‑size rules, drives the observed depth robustness, while a negative result indicates that the stability benefit is limited to non‑normalized layers.

By Fabien Polly
arXiv Machine Learning
1d ago

How Far is Adam from Natural Gradient Descent?

The paper investigates how Adam’s update rule relates to natural gradient descent (NGD) by treating Adam as a diagonal empirical Fisher approximation with additional factors such as diagonal truncation, empirical label substitution, and temporal lag. Using a scale‑invariant metric, the authors quantify Adam’s geometric deviation from true NGD across four loss landscapes—well‑conditioned and ill‑conditioned linear regression, logistic regression, and a small neural network—finding that deviation is low in well‑conditioned settings but can reach about 10³ in ill‑conditioned or non‑convex scenarios. Despite higher geometric drift correlating with slower early optimization, Adam still achieves low final loss, and the improved empirical Fisher (iEF) yields more stable trajectories than the standard empirical Fisher (EF).

By Vihaan Paka-Hegde