arXiv AI

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers

arXiv:2605. 18106v3 Announce Type: replace-cross Abstract: A striking geometric disparity has long persisted in the practice of deep learning.

arXiv Machine Learning
1d ago

CrossGMN: Graph Metanetworks for Cross-Architecture Weight-Space Transformations

CrossGMN introduces a graph metanetwork that processes a trained source network and an initialized target network simultaneously, enabling equivariant cross‑architecture weight‑space transformations. By preserving symmetry through cross‑network message passing, CrossGMN can refine target network initializations while remaining invariant to source permutations and equivariant to target permutations. Experiments demonstrate that CrossGMN accelerates knowledge distillation, transfers across datasets without retraining, and unifies compression from diverse source architectures into a common target architecture.

By Adir Dayan, Yam Eitan, Haggai Maron
arXiv AI
2d ago

GPart: End-to-End Isometric Fine-Tuning via Global Parameter Partitioning

GPart introduces a new parameter‑efficient fine‑tuning technique that directly maps a low‑dimensional trainable vector into the full weight space using a sparse, isometric partition matrix. Unlike LoRA, GPart eliminates the bilinear reconstruction step, preserving exact end‑to‑end isometry and reducing the checkpoint to just the vector and a random seed. Experiments across NLP, vision, and reasoning tasks show that GPart matches or surpasses existing PEFT methods while using far fewer parameters and offering a simpler, more tractable parameterization.

By Paolo Mandica, Micha{\l} Brzozowski, Zuzanna Dubanowska, Neo Christopher Chung
arXiv Machine Learning
1d ago

AF-Muon: An AdamW-Free Muon Optimizer for Tied-Embedding Models

AF‑Muon is an AdamW‑free extension of the Muon optimizer that retains Muon’s matrix update for hidden weights while applying a support‑aware finite‑cap linear minimization oracle to tied vocabulary tables and an RMS‑normalized update for one‑dimensional auxiliary parameters. This design eliminates second‑moment state, reducing optimizer‑state memory by about 20% compared to Hybrid Muon. Across nine tied‑token settings—including decoder‑only language models, T5‑style encoder‑decoders, and ImageGPT‑style variants—AF‑Muon consistently improves mean validation loss and perplexity over both Hybrid Muon and a SCION‑style Sign endpoint, with robust gains confirmed by long‑horizon runs and hyperparameter studies.

By Arash Lagzian, Paniz Halvachi, Junming Zhang, Zhouhan Lin, Dianbo Liu
arXiv Machine Learning
Aug 31

Blog: Survey of Optimizers

The article surveys recent neural‑network optimizers, noting that the field has moved beyond simple Adam variants to encompass matrix‑ and layer‑level designs, time‑policy horizons, and state representations that survive sharding and low‑precision computation. It categorizes optimizers along four axes—temporal estimation, update geometry, horizon management, and representation & systems—highlighting methods such as Muon, Shampoo, SOAP, and quantized states. The survey concludes that while matrix‑aware methods are a genuine advance, no single optimizer universally replaces AdamW, and performance depends on model scale, data‑to‑parameter ratio, batch size, schedule, partitioning, tuning budget, and target metric.

By Ruoran Xu
arXiv Machine Learning
1d ago

TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning

The paper introduces TACO, a new optimizer for fine‑tuning large language models that drastically reduces optimizer state memory while preserving first‑order gradients. TACO selects the sign of the largest magnitude entry in each column of weight matrices, achieving a 174× reduction in persistent optimizer memory compared to AdamW8bit and a 2.9× decrease in peak training memory on OPT‑13B. This allows full‑parameter fine‑tuning of 30–32B‑parameter models on a single 80 GB GPU across multiple model families and tasks, with comparable accuracy and runtime to existing methods.

By Jichao Jiang (University of Central Florida), Cristian McGee (University of Central Florida), El Houcine Bergou (Mohammed VI Polytechnic University), Hanqin Cai (University of Central Florida), Aritra Dutta (University of Central Florida)