arXiv Machine Learning By Gongyue Zhang, Honghai Liu

When the Preconditioning Exponent Turns Negative: Learning-Rate Coupling and Cross-Environment Generalization

Read the original on arXiv Machine Learning →

arXiv:2609. 30271v1 Announce Type: new Abstract: Adaptive optimizers are commonly parameterized by a fixed power of the second-moment estimate.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 7

Optimizer Memory Schedules for Outscaling the Overtraining Axis

The paper studies how different optimizers perform as training duration (overtraining) increases, focusing on matrix‑preconditioned methods (Muon, SOAP) and a momentum‑scheduled method (ADANA) compared to AdamW. Across models ranging from 51M to 253M parameters and overtraining factors up to 256×, the authors find that optimal learning‑rate schedules, weight‑decay coefficients, and memory settings shift with horizon, and that ADANA consistently outperforms AdamW, especially with log‑time weight decay and momentum cooldown. Muon and SOAP maintain roughly constant token‑efficiency advantages, with SOAP potentially improving at the highest overtraining levels.

By Katie Everett, Shikai Qiu