COEC (Calibrated Orthogonal-Equivalence Compensation) is a training‑free framework that improves structured pruning of large language models by applying alternating left and right orthogonal rotations to the retained weight matrix. The method optimizes the right rotation on a reduced Stiefel manifold, rescales singular values via generalized cross‑validation, tempers the calibration Gram matrix, and adds an alignment penalty to preserve geometric relations between attention projections. Experiments on Llama‑3, Llama‑3.1, and Qwen2.5 show that COEC consistently improves perplexity and zero‑shot accuracy across multiple sparsity levels, outperforming existing compensation techniques.
By Peiqi Yu, Nam Ling, Wei Wang, Wei Jiang
arXiv:2609.06557v1 Announce Type: new
Abstract: Large language models (LLMs) are often considered fragile under aggressive sparsification, and maintaining reliable performance typically requires stic...
By Hyeondo Jang, Kwanhee Lee, Dongyeop Lee, Namhoon Lee
arXiv:2605. 18331v2 Announce Type: replace Abstract: Large Language Models (LLMs) have experienced significant growth and development in recent years.
By Diego Coello de Portugal Mecke, Tom Hanika, Lars Schmidth-Thieme
arXiv:2609.00224v1 Announce Type: cross
Abstract: Weight-only post-training quantization (PTQ) can alleviate the computational burden of serving large language models (LLMs) at scale. However, existi...
By Yipin Guo, Arun M George, Jie Fu, Tareq Mahmoud, Sixue Xing, Siddharth Joshi
arXiv:2609.30465v1 Announce Type: cross
Abstract: Mixture-of-experts (MoE) models activate few experts per token but store the full expert pool. Expert pruning reduces this storage burden; at a fixed...
By Mingyang Song, Mao Zheng
Task-Aware Spectral Pruning (TASP) is a post‑training framework that tailors sparse masks to specific tasks by calibrating module‑level spectral descriptors against task‑specific ablation effects. It constructs masks that close grouped‑query‑attention and SwiGLU dependencies, routing each user turn to a single compiled mask that remains fixed during prefill and decoding. In experiments, TASP achieves a 43% active‑FLOP reduction while preserving 97.7% of the dense BF16 performance on Llama‑3‑70B, and delivers a 1.44× speedup on an A100 80GB with INT8‑weight/BF16‑compute, reducing decode latency from 45.2 to 31.3 ms/token.
By Ibne Farabi Shihab, Fariya Afrin, Sanjeda Akter, Anuj Sharma