arXiv AI

Beyond Single-Dimensional Compression: The Compound Sparsity Frontier of Large Language Models

arXiv:2607. 18280v1 Announce Type: cross Abstract: Large language models (LLMs) are often compressed through static parameter pruning or dynamic token-level computation, yet aggressive sparsification can trigger rapid performance degradation beyond an essential sparsity boundary.

arXiv Machine Learning
Aug 27

When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs

Sparse autoencoders (SAEs) are commonly used to interpret large language models, but their reliability after pruning is unclear. This study shows that pruning’s effect on an SAE is governed by perturbation energy, a covariance-weighted norm, and that magnitude pruning distorts the representation space by ignoring activation geometry. Activation-aware pruning methods such as Wanda and SparseGPT better preserve SAE behavior, and the authors find that middle layers are especially vulnerable, leading them to propose a layer‑wise sparsity allocation that reduces perplexity for a given sparsity level.

By Suchit Gupte, Xueru Zhang, Mohammad Mahdi Khalili
arXiv AI
Jul 21

Sparsity-Aware Low-Rank Representation for Efficient Fine-Tuning of Large Language Models

arXiv:2601. 16991v3 Announce Type: replace-cross Abstract: Adapting large pre-trained language models to downstream tasks often entails fine-tuning millions of parameters or deploying costly dense weight updates, which hinders their use in resource-constrained environments.

By Longteng Zhang, Sen Wu, Shuai Hou, Zhengyu Qing, Zhuo Zheng, Danning Ke, Qihong Lin, Qiang Wang, Shaohuai Shi, Xiaowen Chu
arXiv AI
Aug 25

Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling

The paper revisits the impact of pruning on large language models (LLMs) during test-time scaling (TTS). While prior work found that structured pruning degrades reasoning performance, this study shows that unstructured pruning—removing only specific redundant weights—can actually improve TTS performance on reasoning benchmarks for models s1.1-7B and Qwen3-8B, sometimes surpassing the full-weight models. The authors also examine how different layer-wise sparsity allocation strategies affect these outcomes.

By Ocean Monjur, Shahriar Kabir Nahin, Anshuman Chhabra
arXiv Computation and Language
Aug 31

Pruning Laws for Large Language Models

arXiv:2504.04342v2 Announce Type: replace Abstract: Scaling up model parameters and training data consistently improves the performance of large language models (LLMs), but at the cost of rapidly gro...

By Ayan Sengupta, Siddhant Chaudhary, Tanmoy Chakraborty
arXiv AI
Jun 4

L$^3$: Large Lookup Layers

arXiv:2601. 21461v3 Announce Type: replace-cross Abstract: Modern sparse language models typically achieve sparsity through Mixture-of-Experts (MoE) layers, which dynamically route tokens to dense MLP "experts.

By Albert Tseng, Christopher De Sa
arXiv AI
Aug 26

Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression

The paper introduces the "Compression Trinity," a unified framework that jointly applies sparsity, quantization, and low‑rank approximations to compress large language models. It presents several methods—MKOR, SLoPe, OPTIMA, PATCH, and SLiM—that leverage these three pillars to accelerate training, reduce memory bandwidth, and recover accuracy, achieving significant speedups and accuracy gains over existing techniques. The results demonstrate that combining all three compression strategies is essential for efficient, scalable, high‑performance LLM deployment.

By Mohammad Mozaffari
arXiv AI
2d ago

MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs

MWOP (Modality-aware Width-wise Operation Pruning) is a method that independently prunes visual‑to‑visual, text‑to‑visual, and text‑to‑text attention paths within each layer of multimodal large language models, and separately selects feed‑forward network channels for visual and textual inputs. It uses a first‑order Taylor criterion to guide pruning, re‑evaluates FFN importance after attention pruning, and applies LoRA‑based recovery training. The approach is paired with path‑sparse Triton attention kernels and compact visual‑side FFN execution to achieve practical acceleration, preserving token sequences while reducing computation. "whyItMatters":"MWOP achieves a 1.6× prefill speedup on LLaVA‑OneVision‑7B while retaining 99.7% performance, and further boosts token‑compression methods to 2.9× and 2.7× speedups, demonstrating its effectiveness across architectures."

By Xudong Wang, Hao Wu, Haozhe Hu, Peiran Yin, Xinghao Chen, Yunpu Ma, Wei Zhang, Xiaoyu Shen