arXiv Machine Learning By Hu Xu, Zhaolong Xing, Congcong Liu, Jiaxing Wang, Zhida Jiang, Junshi Huang, Zhen Chen, Jianfeng Xu

Averaged Evaluation Masks Capability Trade-Offs: Multi-Source Calibration for High-Sparsity LLM Pruning

Read the original on arXiv Machine Learning →

arXiv:2606. 03328v3 Announce Type: replace Abstract: Calibration data are often treated as a minor implementation detail in post-training LLM pruning because averaged evaluations suggest only modest effects.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 3

Calibration Data Trade-offs Across Capability Dimensions: Why Multi-Source Mixing Matters for High-Sparsity LLM Pruning

arXiv:2606. 03328v1 Announce Type: cross Abstract: Post-training pruning compresses large language models to high sparsity using a small unlabelled calibration set, and recent work has concluded that the choice of calibration source has only modest impact on averaged post-pruning accuracy.

By Hu Xu, Zhaolong Xing, Congcong Liu, Jiaxing Wang, Zhida Jiang, Junshi Huang, Zhen Chen, Jianfeng Xu
arXiv Machine Learning
4d ago

Capability Scaling-Down Laws for LLM Compression

The paper presents a systematic study of how different compression techniques—pruning, quantization, and distillation—affect the capabilities of large language models (LLMs) in tasks such as mathematics, code generation, and question answering. It introduces a framework that measures capability loss and relates it to factors like model size, training stage, and compression settings, yielding simple predictive relations that generalize across unseen configurations. The authors demonstrate that sharing density responses across pruning levels can dramatically reduce the number of measurements needed, and that their predictive models closely match regression results while offering efficient decision guidance for compression method selection.

By Xueqi Cheng, Liang Wu, Kelly Wan, Liangjie Hong, Yushun Dong
arXiv Machine Learning
Sep 25

Task-Aware Spectral Pruning: A Mixture-of-Masks Framework for Efficient LLM Inference

Task-Aware Spectral Pruning (TASP) is a post‑training framework that tailors sparse masks to specific tasks by calibrating module‑level spectral descriptors against task‑specific ablation effects. It constructs masks that close grouped‑query‑attention and SwiGLU dependencies, routing each user turn to a single compiled mask that remains fixed during prefill and decoding. In experiments, TASP achieves a 43% active‑FLOP reduction while preserving 97.7% of the dense BF16 performance on Llama‑3‑70B, and delivers a 1.44× speedup on an A100 80GB with INT8‑weight/BF16‑compute, reducing decode latency from 45.2 to 31.3 ms/token.

By Ibne Farabi Shihab, Fariya Afrin, Sanjeda Akter, Anuj Sharma
arXiv Machine Learning
Jun 17

AIMER: Calibration-Free Task-Agnostic MoE Expert Pruning

arXiv:2603. 18492v3 Announce Type: replace Abstract: Mixture-of-Experts (MoE) language models increase parameter capacity without proportional per-token computation, yet deployment still requires storing the full expert pool, making expert pruning important for reducing memory and serving overhead.

By Zongfang Liu, Guangyi Chen, Shengkun Tang, Yifan Shen, Huan Wang, Xin Yuan
arXiv Computation and Language
Aug 31

Pruning Laws for Large Language Models

arXiv:2504.04342v2 Announce Type: replace Abstract: Scaling up model parameters and training data consistently improves the performance of large language models (LLMs), but at the cost of rapidly gro...

By Ayan Sengupta, Siddhant Chaudhary, Tanmoy Chakraborty