arXiv AI

PowerStep: Memory-Efficient Adaptive Optimization via $\ell_p$-Norm Steepest Descent

arXiv Machine Learning
Sep 25

FlashLoop: Fast and Memory-Efficient Looped Transformers via Lazy Updates

FlashLoop is a training‑free inference framework for Looped Transformers that reduces cross‑loop redundancy by employing token‑sparse updates, sparse attention, and KV‑residual quantization. It exploits observations that, as loops progress, state changes concentrate on a small token subset, attention differences are dominated by a sparse key subset, and KV residuals become amenable to low‑bit quantization. The method achieves lossless accuracy with up to 1.64× speedup and 6× KV‑cache memory reduction across several Looped Transformer models.

By Wanqi Yang, Shiwei Liu
arXiv Machine Learning
Aug 4

AOS: Adaptive Optimizer Switching via Training-State Signals for Faster Convergence and Better Generalization

arXiv:2608. 01997v1 Announce Type: new Abstract: Single-optimizer training is a poor fit for the distinct phases of deep network optimization: adaptive methods handle noisy early gradients well but overshoot flat minima, while SGD with momentum generalizes better in the late phase but converges slowly early on.

By Alok Kumar Pandey, Umang Chaturvedi, Aatish Rana, Gopi Krishna Nedanuri
arXiv Machine Learning
Aug 18

Adaptive Optimization via Momentum on Variance-Normalized Gradients

arXiv:2602. 10204v2 Announce Type: replace Abstract: We introduce MVN-Grad (Momentum on Variance-Normalized Gradients), an Adam-style optimizer that improves stability and performance by combining two complementary ideas: variance-based normalization and momentum applied after normalization.

By Francisco Patitucci, Aryan Mokhtari
arXiv AI
6d ago

G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation

G$^2$PTQ is a post‑training quantization framework that improves large language models by combining first‑ and second‑order information in a globally supervised, block‑wise optimization. It refreshes gradient and Hessian estimates before each Transformer block and uses a trust‑region scaling mechanism to stabilize gradient steps, preventing exploding weight updates. The method achieves better alignment with full‑precision models and outperforms state‑of‑the‑art baselines across various model families and bit‑widths.

By Ruikang Liu, Haoli Bai, Yuxuan Sun, Qian Zhang, Wenzheng Cai, Yanqi Hao, Feiyu Wang, Weidong Zhong, Zhuang Wang, Tong Yang, Xiangsheng Zhou
arXiv AI
Aug 11

Full-bandwidth transformer

arXiv:2608. 08888v1 Announce Type: new Abstract: Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth.

By Xi Wang, Ziyang Cai, Zheng Zhan, Harry Dong, Ying Fan, Gustavo de Rosa, Tim Pearce, John Langford
arXiv Machine Learning
Aug 27

StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models

StoSignSGD is a new sign‑based optimization algorithm that injects structural stochasticity into the sign operator, ensuring unbiased updates. It resolves the divergence issues of traditional SignSGD on non‑smooth objectives, achieving optimal convergence rates in convex settings and improved complexity bounds in non‑convex, non‑smooth problems. Empirical results show that StoSignSGD is stable and efficient across large language model training, outperforming AdamW and SignSGD in low‑precision regimes (FP8 and FP4) and delivering speedups and accuracy gains on models ranging from OLMo2‑370M to 7B LLMs.

By Dingzhi Yu, Rui Pan, Yuxing Liu, Difan Zou, Tong Zhang
arXiv AI
Jun 15

Gefen: Optimized Stochastic Optimizer

arXiv:2606. 13894v1 Announce Type: cross Abstract: AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory.

By Nadav Benedek, Tomer Koren, Ohad Fried
arXiv Machine Learning
Aug 27

Beyond Dense Adam States: Adaptive Log-Space Quantization for Memory-Efficient Optimizers

The paper introduces Adaptive Log‑Space (AL) quantization, a block‑wise representation that adapts the non‑zero range per block and preserves the exact‑zero invariant for non‑negative optimizer states. AL8 and AL16 are combined with signed‑momentum encodings and state‑specific precision choices, rather than a single policy for all states. Experiments on TinyLlama‑1.1B and GPT‑2 show that AL‑based quantization can match or exceed full‑precision performance while dramatically reducing optimizer‑state storage.

By Yan Wang