arXiv AI

Calibration Data Trade-offs Across Capability Dimensions: Why Multi-Source Mixing Matters for High-Sparsity LLM Pruning

arXiv:2606. 03328v1 Announce Type: cross Abstract: Post-training pruning compresses large language models to high sparsity using a small unlabelled calibration set, and recent work has concluded that the choice of calibration source has only modest impact on averaged post-pruning accuracy.

arXiv Machine Learning
4d ago

Capability Scaling-Down Laws for LLM Compression

The paper presents a systematic study of how different compression techniques—pruning, quantization, and distillation—affect the capabilities of large language models (LLMs) in tasks such as mathematics, code generation, and question answering. It introduces a framework that measures capability loss and relates it to factors like model size, training stage, and compression settings, yielding simple predictive relations that generalize across unseen configurations. The authors demonstrate that sharing density responses across pruning levels can dramatically reduce the number of measurements needed, and that their predictive models closely match regression results while offering efficient decision guidance for compression method selection.

By Xueqi Cheng, Liang Wu, Kelly Wan, Liangjie Hong, Yushun Dong
arXiv Machine Learning
Sep 11

LILA: Calibration-Free Structured Pruning of Large Language Models via Latent Spectral Geometry

LILA (Latent-Informed Layer Analysis) introduces a calibration‑free method for structured pruning of large language models by scoring neuron importance using the Kolmogorov–Smirnov distance between singular value distributions of full and neuron‑ablated feed‑forward network weight matrices. The approach requires no training, calibration data, or auxiliary networks, and outperforms existing methods such as PruneNet and SliceGPT on LLaMA‑2‑7B and Phi‑2 at various sparsity levels. After a single epoch of LoRA fine‑tuning, LILA matches heavily calibrated baselines, and a Neural Tangent Kernel analysis provides theoretical support for its spectral importance criterion. Additionally, LILA can dynamically allocate sparsity budgets, achieving state‑of‑the‑art generative preservation and revealing architectural bottlenecks at higher compression.

By Sankar Behera, Dhruv Singh, Anshika Agnihotri, Raj Kumar Choudhary, Satyadev Ahlawat, Yamuna Prasad
arXiv Machine Learning
Jun 17

AIMER: Calibration-Free Task-Agnostic MoE Expert Pruning

arXiv:2603. 18492v3 Announce Type: replace Abstract: Mixture-of-Experts (MoE) language models increase parameter capacity without proportional per-token computation, yet deployment still requires storing the full expert pool, making expert pruning important for reducing memory and serving overhead.

By Zongfang Liu, Guangyi Chen, Shengkun Tang, Yifan Shen, Huan Wang, Xin Yuan
arXiv Machine Learning
Aug 24

COEC: Calibrated Orthogonal-Equivalence Compensation for Structured Pruning of Large Language Models

COEC (Calibrated Orthogonal-Equivalence Compensation) is a training‑free framework that improves structured pruning of large language models by applying alternating left and right orthogonal rotations to the retained weight matrix. The method optimizes the right rotation on a reduced Stiefel manifold, rescales singular values via generalized cross‑validation, tempers the calibration Gram matrix, and adds an alignment penalty to preserve geometric relations between attention projections. Experiments on Llama‑3, Llama‑3.1, and Qwen2.5 show that COEC consistently improves perplexity and zero‑shot accuracy across multiple sparsity levels, outperforming existing compensation techniques.

By Peiqi Yu, Nam Ling, Wei Wang, Wei Jiang
arXiv AI
Aug 28

Frequency Matters: Fast Model-Agnostic Data Curation for Pruning and Quantization

The paper introduces ZipCal, a model‑agnostic data curation method that selects calibration data for post‑training compression of large language models by maximizing lexical diversity using Zipfian power laws. ZipCal outperforms uniform random sampling on pruning benchmarks and matches a state‑of‑the‑art perplexity‑based approach while being roughly 240× faster due to its linear complexity. The authors provide code and experiments at their GitHub repository.

By Francesco Pio Monaco, Elia Cunegatti, Flavio Vella, Giovanni Iacca
arXiv Machine Learning
Sep 25

Task-Aware Spectral Pruning: A Mixture-of-Masks Framework for Efficient LLM Inference

Task-Aware Spectral Pruning (TASP) is a post‑training framework that tailors sparse masks to specific tasks by calibrating module‑level spectral descriptors against task‑specific ablation effects. It constructs masks that close grouped‑query‑attention and SwiGLU dependencies, routing each user turn to a single compiled mask that remains fixed during prefill and decoding. In experiments, TASP achieves a 43% active‑FLOP reduction while preserving 97.7% of the dense BF16 performance on Llama‑3‑70B, and delivers a 1.44× speedup on an A100 80GB with INT8‑weight/BF16‑compute, reducing decode latency from 45.2 to 31.3 ms/token.

By Ibne Farabi Shihab, Fariya Afrin, Sanjeda Akter, Anuj Sharma
arXiv AI
Sep 17

Higher-order pruning of experts in mixture-of-experts language models

The paper introduces HOPE, a second‑order pruning method for Mixture‑of‑Experts language models that accounts for cooperative interactions between experts. Unlike first‑order methods such as REAP, HOPE derives an objective that provably bounds pruning error and is shown to outperform baselines across three large MoE models, multiple calibration sets, and diverse benchmarks, especially at high pruning rates and on agentic tasks. The results demonstrate that preserving expert interactions allows aggressive compression with minimal performance loss on complex workloads.

By Alex M. Tseng, Prannay Kaul, Luca Zancato, Wei Xia, Stefano Soatto