arXiv Machine Learning By Felix Stollenwerk, Anna Lokrantz, Niclas Hertzberg

Output Embedding Centering for Stable LLM Pretraining

Read the original on arXiv Machine Learning →

The paper addresses training instabilities in large language model pretraining, specifically output logit divergence that occurs near the end of training. By analyzing the geometry of output embeddings, the authors identify anisotropic embeddings as the root cause and propose Output Embedding Centering (OEC) as a mitigation strategy. OEC can be applied deterministically as μ‑centering or as a regularization loss μ‑loss, and experiments show both variants outperform the existing z‑loss method while matching logit soft‑capping in stability, even without weight tying. Additionally, μ‑loss is less sensitive to hyperparameter tuning than z‑loss.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 5

Dominant-Layer ZO: A Single Layer Dominates Zeroth-Order Fine-Tuning of LLMs

arXiv:2606. 05516v1 Announce Type: new Abstract: Zeroth-order (ZO) optimization enables memory-efficient fine-tuning of large language models (LLMs) using only forward passes, but it remains unclear how useful adaptation is distributed across layers.

By Wanhao Yu, Ziyan Wang, Zheng Wang, Abeer Matar Almalky, Yihang Zuo, Shuteng Niu, Sen Lin, Adnan Siraj Rakin, Deliang Fan, Li Yang
arXiv Machine Learning
Sep 4

MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration

The paper introduces MSign, an optimizer designed to prevent training instability in large language models by restoring the stable rank of weight matrices. It identifies two precursors to gradient explosions—rapid stable rank decline and increased Jacobian alignment—and proves that these jointly cause exponential gradient growth. Experiments on models ranging from 5 M to 3 B parameters show that MSign stops training failures while adding less than 7.0% computational overhead.

By Lianhai Ren, Yucheng Ding, Xiao Liu, Peng Cheng, Yeyun Gong