The Sparsity Whisperer
arXiv:2608. 06630v1 Announce Type: new Abstract: Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs.
The paper introduces a layerwise, decoupled approach to structurally sparsify fully connected layers in pretrained neural networks. By extracting shallow two‑layer subnetworks, normalizing inner weights, and applying a structured group penalty to each block’s outer weight matrix, the method prunes neurons sequentially and reduces layer widths. The authors prove equivalence to a joint penalty for positively homogeneous activations, and demonstrate that this decoupled formulation is more robust, offering a broader regularization range and lower catastrophic over‑pruning while preserving accuracy in classification, sparse‑recovery, PINN, and OPT‑1.3B experiments.
arXiv:2608. 06630v1 Announce Type: new Abstract: Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs.
arXiv:2606. 14346v1 Announce Type: cross Abstract: Unstructured pruning produces sparse weight tensors, but the standard implementation keeps tensor shapes unchanged so the deployed model is no smaller than before pruning.
arXiv:2608. 16010v1 Announce Type: new Abstract: Model compression is critical for deploying networks on resource-constrained edge devices.
arXiv:2607. 21366v1 Announce Type: cross Abstract: Deep neural networks encode complex representations, but deconstructing this internal knowledge remains a challenge.
arXiv:2607. 20555v1 Announce Type: new Abstract: The lottery ticket hypothesis proposes that large random neural networks contain sparse subnetworks that can match the performance of dense models after comparable training.
arXiv:2609.24401v1 Announce Type: new Abstract: Structured pruning is a model compression technique that is used to reduce the computational cost of deploying deep neural networks on resource-constra...
arXiv:2606. 30676v1 Announce Type: cross Abstract: Deploying spiking neural networks (SNNs) on neuromorphic hardware demands aggressive synaptic pruning while preserving temporal computation integrity.
arXiv:2609.40127v1 Announce Type: cross Abstract: Modern transformers pair impressive capabilities with substantial memory and compute demands. Low-rank weight factorization reduces both while keepin...
arXiv:2606. 27759v1 Announce Type: new Abstract: Training binary neural networks (BNNs) from scratch is dominated by the straight-through estimator (STE), whose forward/backward mismatch produces severe accuracy degradation as networks deepen.
arXiv:2606. 09885v1 Announce Type: new Abstract: Mixture-of-Experts large language models (LLMs) scale efficiently through sparse activation, yet their deployment is fundamentally constrained by the large static parameter footprint of experts.
Neural network pruning reduces model size by removing less important parameters while aiming to preserve predictive performance. Although the Lottery Ticket Hypothesis (LTH) shows that sparse subnetworks can match dense networks when trained from suitable initializations, its iterative pruning procedure requires multiple complete training cycles.
arXiv:2606. 12278v1 Announce Type: cross Abstract: Neural network pruning reduces model size by removing less important parameters while aiming to preserve predictive performance.