arXiv Machine Learning

Z-Loss Backward Geometry in Dense Output Heads and Sparse Routers

arXiv Machine Learning
Jun 5

Dominant-Layer ZO: A Single Layer Dominates Zeroth-Order Fine-Tuning of LLMs

arXiv:2606. 05516v1 Announce Type: new Abstract: Zeroth-order (ZO) optimization enables memory-efficient fine-tuning of large language models (LLMs) using only forward passes, but it remains unclear how useful adaptation is distributed across layers.

By Wanhao Yu, Ziyan Wang, Zheng Wang, Abeer Matar Almalky, Yihang Zuo, Shuteng Niu, Sen Lin, Adnan Siraj Rakin, Deliang Fan, Li Yang
arXiv Machine Learning
1d ago

Scaling Zero-Order Pretraining through Model Sharding

arXiv:2609.37899v1 Announce Type: new Abstract: Zero-order optimization (ZO) trains without backpropagation, making it relevant to forward-only hardware and non-differentiable loss, but its gradient...

By Francois Chaubard, Mykel J. Kochenderfer, Chris R\'e
arXiv Computer Vision
Aug 27

SHIFT-LLM: Distribution Shift Correction in Depth-Pruned LLMs

SHIFT-LLM is a training‑free post‑pruning correction framework that inserts a Linear Residual Adapter (LRA) at each depth‑pruned site in large language models. Each LRA preserves the original residual identity while adding a lightweight affine correction calibrated via closed‑form least‑squares regression on a small held‑out set, thereby approximating the hidden state that would have been produced by the removed block. Experiments across multiple model families and benchmarks show that SHIFT‑LLM consistently recovers accuracy lost to depth pruning, achieving gains up to +15.7 points on Llama‑3.1‑8B‑Instruct with only a few hundred calibration samples and no gradient computation.

By Ali Bahri, Hang Li, Hongliang Li, Zhitang Chen
arXiv Machine Learning
Sep 11

Output Embedding Centering for Stable LLM Pretraining

The paper addresses training instabilities in large language model pretraining, specifically output logit divergence that occurs near the end of training. By analyzing the geometry of output embeddings, the authors identify anisotropic embeddings as the root cause and propose Output Embedding Centering (OEC) as a mitigation strategy. OEC can be applied deterministically as μ‑centering or as a regularization loss μ‑loss, and experiments show both variants outperform the existing z‑loss method while matching logit soft‑capping in stability, even without weight tying. Additionally, μ‑loss is less sensitive to hyperparameter tuning than z‑loss.

By Felix Stollenwerk, Anna Lokrantz, Niclas Hertzberg
arXiv Computation and Language
6d ago

Baseline Shape Decides the Verdict: A Controlled Re-Examination of Ternary Language Models at 60K Parameters

The study re‑examines a reported advantage of a routed ternary (1.58‑bit) language model over a full‑precision transformer at 60K parameters. By running controlled experiments with multiple seeds and a fixed training recipe, the authors find that the apparent benefit largely stems from the choice of baseline model shape rather than the ternary architecture itself. While the routed model does outperform other shapes at a larger 130M‑byte budget, its advantage diminishes when a plain gated diagonal‑SSM block is used, and the ternary penalty varies with architecture and quantization details.

By Gautam Veldanda
arXiv AI
Sep 15

LayerRoute: Adaptive Layer-Skipping with LoRA-Preserved Quality for Efficient LLM Inference

LayerRoute is a parameter‑efficient technique that enables adaptive skipping of transformer layers in large language models. It adds a lightweight per‑layer router (~21.5K parameters) and LoRA adapters (rank 8, ~1.08M parameters) to each of the 24 blocks in Qwen2.5‑0.5B‑Instruct, training them jointly with a gate‑regularized language‑modeling objective. Across ten independent runs, the method consistently identifies nine middle layers as skip‑eligible, achieves a verified wall‑clock speedup of 1.02x–1.06x, and improves perplexity by an average of 1.16 points, while the router’s decisions vary per input, confirming genuine adaptive behavior.

By Prateek Kumar Sikdar
arXiv AI
Jun 17

Rethinking Cross-Layer Information Routing in Diffusion Transformers

arXiv:2605. 20708v2 Announce Type: replace-cross Abstract: Diffusion Transformers (DiTs) have become a de facto backbone of modern visual generation, and nearly every major axis of their design -- tokenization, attention, conditioning, objectives, and latent autoencoders -- has been extensively revisited.

By Chao Xu, Maohua Li, Qirui Li, Yixuan Xu, Yanke Zhou, Yunhe Li, Cuifeng Shen, Hanlin Tang, Kan Liu, Tao Lan, Lin Qu, Shao-Qun Zhang