arXiv Machine Learning By Da Chang, Ganzhao Yuan

MGUP: A Momentum-Gradient Alignment Update Policy for Stochastic Optimization

Read the original on arXiv Machine Learning →

arXiv:2606. 17526v1 Announce Type: new Abstract: Efficient optimization is essential for training large language models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 18

Adaptive Optimization via Momentum on Variance-Normalized Gradients

arXiv:2602. 10204v2 Announce Type: replace Abstract: We introduce MVN-Grad (Momentum on Variance-Normalized Gradients), an Adam-style optimizer that improves stability and performance by combining two complementary ideas: variance-based normalization and momentum applied after normalization.

By Francisco Patitucci, Aryan Mokhtari
arXiv Machine Learning
Aug 27

StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models

StoSignSGD is a new sign‑based optimization algorithm that injects structural stochasticity into the sign operator, ensuring unbiased updates. It resolves the divergence issues of traditional SignSGD on non‑smooth objectives, achieving optimal convergence rates in convex settings and improved complexity bounds in non‑convex, non‑smooth problems. Empirical results show that StoSignSGD is stable and efficient across large language model training, outperforming AdamW and SignSGD in low‑precision regimes (FP8 and FP4) and delivering speedups and accuracy gains on models ranging from OLMo2‑370M to 7B LLMs.

By Dingzhi Yu, Rui Pan, Yuxing Liu, Difan Zou, Tong Zhang
arXiv AI
Jun 15

Gefen: Optimized Stochastic Optimizer

arXiv:2606. 13894v1 Announce Type: cross Abstract: AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory.

By Nadav Benedek, Tomer Koren, Ohad Fried