arXiv Machine Learning

How Bregman Divergences Shape Shampoo

arXiv Machine Learning
Jun 26

DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers

arXiv:2602. 02016v2 Announce Type: replace Abstract: Shampoo is one of the leading approximate second-order optimizers: a variant of it has won the MLCommons AlgoPerf competition, and it has been shown to produce models with lower activation outliers that are easier to compress.

By Ionut-Vlad Modoranu, Philip Zmushko, Erik Schultheis, Mher Safaryan, Dan Alistarh
arXiv Machine Learning
Aug 31

Blog: Survey of Optimizers

The article surveys recent neural‑network optimizers, noting that the field has moved beyond simple Adam variants to encompass matrix‑ and layer‑level designs, time‑policy horizons, and state representations that survive sharding and low‑precision computation. It categorizes optimizers along four axes—temporal estimation, update geometry, horizon management, and representation & systems—highlighting methods such as Muon, Shampoo, SOAP, and quantized states. The survey concludes that while matrix‑aware methods are a genuine advance, no single optimizer universally replaces AdamW, and performance depends on model scale, data‑to‑parameter ratio, batch size, schedule, partitioning, tuning budget, and target metric.

By Ruoran Xu
arXiv Statistics ML
6d ago

Exact information accounting for SGD methods

The paper presents an exact, information‑theoretic analysis of stochastic gradient descent (SGD) and its variants, showing that a preconditioned SGD step corresponds to a posterior‑mean update in a Gaussian Bayes model. It decomposes one‑step regret into an intrinsic‑time cost and a change in comparator information, extending this split to an identity for the objective itself. The framework links convex convergence, saddle‑point escape, flatness‑generalization trade‑offs, learning‑rate schedules, adaptive optimizers, and various SGD variants, and it is validated on synthetic and real training runs, revealing how different optimizers achieve the same training loss through distinct step characteristics.

By Akshay Balsubramani
arXiv Machine Learning
Sep 7

Optimizer Memory Schedules for Outscaling the Overtraining Axis

The paper studies how different optimizers perform as training duration (overtraining) increases, focusing on matrix‑preconditioned methods (Muon, SOAP) and a momentum‑scheduled method (ADANA) compared to AdamW. Across models ranging from 51M to 253M parameters and overtraining factors up to 256×, the authors find that optimal learning‑rate schedules, weight‑decay coefficients, and memory settings shift with horizon, and that ADANA consistently outperforms AdamW, especially with log‑time weight decay and momentum cooldown. Muon and SOAP maintain roughly constant token‑efficiency advantages, with SOAP potentially improving at the highest overtraining levels.

By Katie Everett, Shikai Qiu