arXiv Machine Learning

AutoNorm: Understanding Adaptive Normalization in Transformers through Differentiable Gating

arXiv:2607. 10593v1 Announce Type: new Abstract: Normalization is a critical component for stabilizing Transformer training, yet the choice between static strategies such as Layer Normalization (LN) and adaptive alternatives remains largely task-dependent.

arXiv AI
Aug 19

Gradient Heterogeneity Complements Hessian Heterogeneity in Transformer Optimization

The paper investigates why adaptive optimizers like Adam outperform SGD when fine‑tuning Transformers. It introduces gradient heterogeneity—the variation in gradient norms across parameter blocks—and shows, both theoretically and experimentally, that this heterogeneity, together with Hessian heterogeneity, hampers SGD convergence while sign‑based methods such as SignSGD are less affected. The study links the source of gradient heterogeneity to layer‑normalization placement, finding that Post‑LN architectures exhibit the strongest effect, and uses SignSGD as a tractable proxy to analyze Adam‑like behavior and learning‑rate scaling.

By Akiyoshi Tomihari, Issei Sato
Hugging Face Trending Papers
Aug 9

Gradient Under Microscope: Benchmarking Resource Utilization of Memory-Efficient Gradient Computation Methods

AI training's rising resource intensity is straining electricity supplies and carbon budgets, motivating systematic study of memory-efficient training on constrained hardware. We benchmark five gradient optimizers (SGD, Adam, Adagrad, Adadelta, and Conjugate Gradient Descent) under three memory strategies (standard training, gradient checkpointing, and gradient accumulation) across four transformer architectures (ViT, ModernBERT, Llama 3.

arXiv AI
Sep 24

Anon: Extrapolating Adaptivity Beyond SGD and Adam

The paper introduces Anon, an optimizer that extends adaptivity beyond the traditional bounds of SGD and Adam by allowing extrapolation across the entire real-number spectrum. It addresses the limitations of prior tunable optimizers that only interpolate between 0 and 1 adaptivity, showing that optimal adaptivity can require negative values for CNNs or values greater than one for Transformers. Anon incorporates Incremental Delay Update (IDU) to maintain provable stability and demonstrates competitive performance on image classification, diffusion, and large language modeling tasks.

By Yiheng Zhang, Kaiyan Zhao, Shaowu Wu, Yiming Wang, Jiajun Wu, Leong Hou U, Steve Drew, Xiaoguang Niu
arXiv Machine Learning
Sep 10

Conditioned Initialization for Attention

arXiv:2609.07086v1 Announce Type: new Abstract: Transformers are a dominant architecture in modern machine learning, powering applications across vision, language, and beyond. At the core of their su...

By Hemanth Saratchandran, Simon Lucey
arXiv Machine Learning
Sep 11

SG-Blend: Learning an Interpolation Between Improved Swish and GELU for Robust Neural Representations

SG-Blend introduces a per‑layer adaptive activation that interpolates between a bias‑corrected, parametric Swish variant (SSwish) and GELU, using a learnable blend coefficient, sharpness, and zero‑centering bias. The method adds only three scalars per feed‑forward block and, on BERT‑style IMDB classification, matches peak accuracy while reducing seed‑to‑seed variance by 42 %. It also achieves the lowest validation perplexity on WikiText103 and generalizes to computer vision and other domains.

By Gaurav Sarkar, Syed Affan Daimi, Jay Gala, Subarna Tripathi