arXiv Machine Learning By Jichao Jiang (University of Central Florida), Cristian McGee (University of Central Florida), El Houcine Bergou (Mohammed VI Polytechnic University), Hanqin Cai (University of Central Florida), Aritra Dutta (University of Central Florida)

TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning

Read the original on arXiv Machine Learning →

The paper introduces TACO, a new optimizer for fine‑tuning large language models that drastically reduces optimizer state memory while preserving first‑order gradients. TACO selects the sign of the largest magnitude entry in each column of weight matrices, achieving a 174× reduction in persistent optimizer memory compared to AdamW8bit and a 2.9× decrease in peak training memory on OPT‑13B. This allows full‑parameter fine‑tuning of 30–32B‑parameter models on a single 80 GB GPU across multiple model families and tasks, with comparable accuracy and runtime to existing methods.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 12

MiniMax Sparse Attention

arXiv:2606. 13392v1 Announce Type: new Abstract: Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persistent memory all require the model to jointly attend over hundreds of thousands to millions of tokens, yet the quadratic cost of softmax attention makes this untenable at deployment scale.

By Xunhao Lai, Weiqi Xu, Yufeng Yang, Qiaorui Chen, Yang Xu, Lunbin Zeng, Xiaolong Li, Haohai Sun, Haichao Zhu, Vito Zhang, Pengyu Zhao
arXiv Machine Learning
1d ago

AF-Muon: An AdamW-Free Muon Optimizer for Tied-Embedding Models

AF‑Muon is an AdamW‑free extension of the Muon optimizer that retains Muon’s matrix update for hidden weights while applying a support‑aware finite‑cap linear minimization oracle to tied vocabulary tables and an RMS‑normalized update for one‑dimensional auxiliary parameters. This design eliminates second‑moment state, reducing optimizer‑state memory by about 20% compared to Hybrid Muon. Across nine tied‑token settings—including decoder‑only language models, T5‑style encoder‑decoders, and ImageGPT‑style variants—AF‑Muon consistently improves mean validation loss and perplexity over both Hybrid Muon and a SCION‑style Sign endpoint, with robust gains confirmed by long‑horizon runs and hyperparameter studies.

By Arash Lagzian, Paniz Halvachi, Junming Zhang, Zhouhan Lin, Dianbo Liu
arXiv AI
Jun 15

Gefen: Optimized Stochastic Optimizer

arXiv:2606. 13894v1 Announce Type: cross Abstract: AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory.

By Nadav Benedek, Tomer Koren, Ohad Fried
arXiv Machine Learning
Sep 23

MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training

MONA is a new optimizer that extends the Muon optimizer by adding a Nesterov‑style acceleration term derived from an exponential moving average of gradient differences. The paper provides a convergence analysis showing that this term offers curvature‑aware corrections while maintaining Muon’s spectral‑norm regularization. Empirical results demonstrate that MONA outperforms both Muon and AdamW on Mixture‑of‑Experts pretraining across models ranging from 1 B to 68 B parameters, and achieves state‑of‑the‑art performance on downstream benchmarks after fine‑tuning the largest model.

By Jiacheng Li, Jianchao Tan, Hongtao Xu, Jiaqi Zhang, Yifan Lu, Yerui Sun, Yuchen Xie, Xunliang Cai