arXiv Machine Learning

DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers

arXiv:2602. 02016v2 Announce Type: replace Abstract: Shampoo is one of the leading approximate second-order optimizers: a variant of it has won the MLCommons AlgoPerf competition, and it has been shown to produce models with lower activation outliers that are easier to compress.

arXiv Machine Learning
1d ago

How Bregman Divergences Shape Shampoo

arXiv:2610.08534v1 Announce Type: new Abstract: Understanding the principles behind Shampoo has recently guided the development of more effective neural network optimizers. These methods learn a prec...

By Bing Liu, Wenjie Zhou, Chengcheng Zhao, Hongtao Zhang, Boao Kong, Felix Dangel, Wu Lin
arXiv Machine Learning
Aug 31

Blog: Survey of Optimizers

The article surveys recent neural‑network optimizers, noting that the field has moved beyond simple Adam variants to encompass matrix‑ and layer‑level designs, time‑policy horizons, and state representations that survive sharding and low‑precision computation. It categorizes optimizers along four axes—temporal estimation, update geometry, horizon management, and representation & systems—highlighting methods such as Muon, Shampoo, SOAP, and quantized states. The survey concludes that while matrix‑aware methods are a genuine advance, no single optimizer universally replaces AdamW, and performance depends on model scale, data‑to‑parameter ratio, batch size, schedule, partitioning, tuning budget, and target metric.

By Ruoran Xu
arXiv Machine Learning
Sep 4

Efficient Constant Optimization for Symbolic Regression with GPU-Accelerated Tree-Based Genetic Programming

The paper introduces a GPU-resident, batched Levenberg–Marquardt solver that efficiently optimizes constants in tree-based genetic programming for symbolic regression. By using reverse-mode automatic differentiation to assemble per-tree Jacobians in a single backward sweep, the solver’s per-iteration cost becomes independent of the number of constants per tree, achieving up to 510,000 trees per second on an NVIDIA A100. Integrated into EvoGP, the solver enables end-to-end search that recovers governing equations on 10 of 18 constructed problems, a significant improvement over stock EvoGP.

By Hao Mao, Xu Tony Liu, Shuai Lu, Peng Zhao, Wenzheng Jiang, Yuntian Chen
arXiv AI
Jul 24

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

arXiv:2607. 20548v1 Announce Type: cross Abstract: Higher-order optimizers such as Muon and SOAP offer faster convergence than AdamW, but their computational cost and numerical stability challenges have limited adoption at scale.

By Mikail Khona, Aditya Vavre, Boxiang Wang, Deyu Fu, Hao Wu, Mike Chrzanowski, Bryan Catanzaro, Dheevatsa Mudigere, Jeff Pool, Michael Lightstone, Mohammad Shoeybi, Mostofa Patwary, Nima Tajbakhsh, Tijmen Blankevoort
arXiv Machine Learning
Sep 3

GRADSOLVE: fast exact gradients for ODE ensembles on GPUs

GRADSOLVE is an open‑source JAX library that provides fast, exact reverse‑mode gradients for low‑dimensional ordinary differential equation (ODE) ensembles on NVIDIA GPUs. It records the accepted steps of an adaptive solver and differentiates a fixed‑step replay, yielding the exact discrete adjoint at a lower computational cost than traditional checkpointed methods. Benchmarks show that GRADSOLVE’s forward kernel is 2.8× faster than DiffEqGPU.jl, and its gradient computation is 5.6–14.1× faster than Diffrax’s checkpointed adjoint while maintaining matched forward‑state accuracy across multiple GPU generations.

By Alessio Spurio Mancini