arXiv Machine Learning

Layerwise Decoupling for Stable Structured Sparsification of Fully Connected Layers

The paper introduces a layerwise, decoupled approach to structurally sparsify fully connected layers in pretrained neural networks. By extracting shallow two‑layer subnetworks, normalizing inner weights, and applying a structured group penalty to each block’s outer weight matrix, the method prunes neurons sequentially and reduces layer widths. The authors prove equivalence to a joint penalty for positively homogeneous activations, and demonstrate that this decoupled formulation is more robust, offering a broader regularization range and lower catastrophic over‑pruning while preserving accuracy in classification, sparse‑recovery, PINN, and OPT‑1.3B experiments.

arXiv Machine Learning
Aug 10

The Sparsity Whisperer

arXiv:2608. 06630v1 Announce Type: new Abstract: Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs.

By Linghao Kong, Inimai Subramanian, Micah Adler, Dan Alistarh, Dan Gutfreund, Nir Shavit
arXiv Computation and Language
3d ago

Learning Functional Subspaces for Neural Network Compression

arXiv:2609.40127v1 Announce Type: cross Abstract: Modern transformers pair impressive capabilities with substantial memory and compute demands. Low-rank weight factorization reduces both while keepin...

By Massimo Bini, Anders Christensen, Stephan Alaniz, Judah Goldfeder, Ole Winther, Yann LeCun, Ravid Shwartz-Ziv, Zeynep Akata
Hugging Face Trending Papers
Jun 10

Finding Sparse Subnetworks in One Training Cycle via Progressive Magnitude-Based Pruning

Neural network pruning reduces model size by removing less important parameters while aiming to preserve predictive performance. Although the Lottery Ticket Hypothesis (LTH) shows that sparse subnetworks can match dense networks when trained from suitable initializations, its iterative pruning procedure requires multiple complete training cycles.