arXiv:2606. 03328v1 Announce Type: cross Abstract: Post-training pruning compresses large language models to high sparsity using a small unlabelled calibration set, and recent work has concluded that the choice of calibration source has only modest impact on averaged post-pruning accuracy.
By Hu Xu, Zhaolong Xing, Congcong Liu, Jiaxing Wang, Zhida Jiang, Junshi Huang, Zhen Chen, Jianfeng Xu
The paper presents a systematic study of how different compression techniques—pruning, quantization, and distillation—affect the capabilities of large language models (LLMs) in tasks such as mathematics, code generation, and question answering. It introduces a framework that measures capability loss and relates it to factors like model size, training stage, and compression settings, yielding simple predictive relations that generalize across unseen configurations. The authors demonstrate that sharing density responses across pruning levels can dramatically reduce the number of measurements needed, and that their predictive models closely match regression results while offering efficient decision guidance for compression method selection.
By Xueqi Cheng, Liang Wu, Kelly Wan, Liangjie Hong, Yushun Dong
Task-Aware Spectral Pruning (TASP) is a post‑training framework that tailors sparse masks to specific tasks by calibrating module‑level spectral descriptors against task‑specific ablation effects. It constructs masks that close grouped‑query‑attention and SwiGLU dependencies, routing each user turn to a single compiled mask that remains fixed during prefill and decoding. In experiments, TASP achieves a 43% active‑FLOP reduction while preserving 97.7% of the dense BF16 performance on Llama‑3‑70B, and delivers a 1.44× speedup on an A100 80GB with INT8‑weight/BF16‑compute, reducing decode latency from 45.2 to 31.3 ms/token.
By Ibne Farabi Shihab, Fariya Afrin, Sanjeda Akter, Anuj Sharma
arXiv:2603. 18492v3 Announce Type: replace Abstract: Mixture-of-Experts (MoE) language models increase parameter capacity without proportional per-token computation, yet deployment still requires storing the full expert pool, making expert pruning important for reducing memory and serving overhead.
By Zongfang Liu, Guangyi Chen, Shengkun Tang, Yifan Shen, Huan Wang, Xin Yuan
arXiv:2607. 16721v1 Announce Type: new Abstract: The strongest open-weight coding models are mixture-of-experts (MoE) networks: most of their size comes from large pools of "expert" subnetworks, of which only a few act on any token.
By Anik Jha
arXiv:2504.04342v2 Announce Type: replace
Abstract: Scaling up model parameters and training data consistently improves the performance of large language models (LLMs), but at the cost of rapidly gro...
By Ayan Sengupta, Siddhant Chaudhary, Tanmoy Chakraborty
arXiv:2608.23744v1 Announce Type: new
Abstract: Split conformal prediction, not the pruning rule, supplies finite-sample marginal coverage once a pruned model is fixed independently of the conformal...
By Ibne Farabi Shihab, Adria Binte Habib, Anuj Sharma
arXiv:2609.21208v1 Announce Type: new
Abstract: Self-play methods that co-train a single language model as both coder and test author promise to move code-generation RL beyond fixed test suites, but...
By Ana Nunez, Peyman Najafirad
arXiv:2608. 12953v1 Announce Type: cross Abstract: Structured pruning is a promising approach for compressing large language models (LLMs), yet existing methods rely heavily on greedy heuristics that produce myopic decisions, and often fail to precisely meet target compression budgets.
By Palaash Goel, Ayan Sengupta, Akshay Nambi, Tanmoy Chakraborty
COEC (Calibrated Orthogonal-Equivalence Compensation) is a training‑free framework that improves structured pruning of large language models by applying alternating left and right orthogonal rotations to the retained weight matrix. The method optimizes the right rotation on a reduced Stiefel manifold, rescales singular values via generalized cross‑validation, tempers the calibration Gram matrix, and adds an alignment penalty to preserve geometric relations between attention projections. Experiments on Llama‑3, Llama‑3.1, and Qwen2.5 show that COEC consistently improves perplexity and zero‑shot accuracy across multiple sparsity levels, outperforming existing compensation techniques.
By Peiqi Yu, Nam Ling, Wei Wang, Wei Jiang
The paper introduces HOPE, a second‑order pruning method for Mixture‑of‑Experts language models that accounts for cooperative interactions between experts. Unlike first‑order methods such as REAP, HOPE derives an objective that provably bounds pruning error and is shown to outperform baselines across three large MoE models, multiple calibration sets, and diverse benchmarks, especially at high pruning rates and on agentic tasks. The results demonstrate that preserving expert interactions allows aggressive compression with minimal performance loss on complex workloads.
By Alex M. Tseng, Prannay Kaul, Luca Zancato, Wei Xia, Stefano Soatto
arXiv:2609.25809v1 Announce Type: new
Abstract: Fine-grained mixture-of-experts (MoE) architectures have become a mainstream design for open-weight LLMs, with hundreds of experts and increasingly man...
By Yuanteng Chen, Qiwei Lai, Chen Tianqi, Peisong Wang, Yuantian Shao, Nanxin Zeng, Zhilei Liu, Chuangyi Li, Jing Liu, Jian Cheng