arXiv:2601. 22580v2 Announce Type: replace-cross Abstract: The success of Large Language Models (LLMs) hinges on the stable training of deep Transformer architectures.
By Chao Wang, Bei Li, Jiaqi Zhang, Xinyu Liu, Yuchun Fan, Linkun Lyu, Xin Chen, Jingang Wang, Tong Xiao, Peng Pei, Xunliang Cai
arXiv:2410. 24050v3 Announce Type: replace Abstract: Large-scale pretraining of transformers has been central to the success of foundation models.
By Ambroise Odonnat, Wassim Bouaziz, Vivien Cabannes
The paper investigates why adaptive optimizers like Adam outperform SGD when fine‑tuning Transformers. It introduces gradient heterogeneity—the variation in gradient norms across parameter blocks—and shows, both theoretically and experimentally, that this heterogeneity, together with Hessian heterogeneity, hampers SGD convergence while sign‑based methods such as SignSGD are less affected. The study links the source of gradient heterogeneity to layer‑normalization placement, finding that Post‑LN architectures exhibit the strongest effect, and uses SignSGD as a tractable proxy to analyze Adam‑like behavior and learning‑rate scaling.
By Akiyoshi Tomihari, Issei Sato
arXiv:2510. 09904v2 Announce Type: replace-cross Abstract: Despite their widespread use, training deep Transformers can be unstable.
By Kelvin Kan, Xingjian Li, Benjamin J. Zhang, Tuhin Sahai, Stanley Osher, Krishna Kumar, Markos A. Katsoulakis
AI training's rising resource intensity is straining electricity supplies and carbon budgets, motivating systematic study of memory-efficient training on constrained hardware. We benchmark five gradient optimizers (SGD, Adam, Adagrad, Adadelta, and Conjugate Gradient Descent) under three memory strategies (standard training, gradient checkpointing, and gradient accumulation) across four transformer architectures (ViT, ModernBERT, Llama 3.
The paper introduces Anon, an optimizer that extends adaptivity beyond the traditional bounds of SGD and Adam by allowing extrapolation across the entire real-number spectrum. It addresses the limitations of prior tunable optimizers that only interpolate between 0 and 1 adaptivity, showing that optimal adaptivity can require negative values for CNNs or values greater than one for Transformers. Anon incorporates Incremental Delay Update (IDU) to maintain provable stability and demonstrates competitive performance on image classification, diffusion, and large language modeling tasks.
By Yiheng Zhang, Kaiyan Zhao, Shaowu Wu, Yiming Wang, Jiajun Wu, Leong Hou U, Steve Drew, Xiaoguang Niu