arXiv Machine Learning By Jundong Hu, Shekar Ramachandran

XMerge: Cross-Axis Selection and Reconstructive Layer Merging for LLM Depth Compression

Read the original on arXiv Machine Learning →

XMerge is a post‑training method for compressing large language models by removing entire transformer layers while preserving a standard serving architecture. It selects low‑impact blocks via cross‑axis selection and refits adjacent surviving blocks with local boundary reconstruction, requiring no task labels or fine‑tuning. Across seven Llama and Qwen backbones, XMerge outperforms five published baselines, achieving top rankings on CORE and MMLU tasks even at aggressive compression levels, and consistently avoids model collapse while improving calibration and decoding efficiency.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
1d ago

Learning Functional Subspaces for Neural Network Compression

arXiv:2609.40127v1 Announce Type: cross Abstract: Modern transformers pair impressive capabilities with substantial memory and compute demands. Low-rank weight factorization reduces both while keepin...

By Massimo Bini, Anders Christensen, Stephan Alaniz, Judah Goldfeder, Ole Winther, Yann LeCun, Ravid Shwartz-Ziv, Zeynep Akata
arXiv Computer Vision
Aug 27

SHIFT-LLM: Distribution Shift Correction in Depth-Pruned LLMs

SHIFT-LLM is a training‑free post‑pruning correction framework that inserts a Linear Residual Adapter (LRA) at each depth‑pruned site in large language models. Each LRA preserves the original residual identity while adding a lightweight affine correction calibrated via closed‑form least‑squares regression on a small held‑out set, thereby approximating the hidden state that would have been produced by the removed block. Experiments across multiple model families and benchmarks show that SHIFT‑LLM consistently recovers accuracy lost to depth pruning, achieving gains up to +15.7 points on Llama‑3.1‑8B‑Instruct with only a few hundred calibration samples and no gradient computation.

By Ali Bahri, Hang Li, Hongliang Li, Zhitang Chen