arXiv Machine Learning

Double-Scoring: Reliable Extraction of Strong Lottery Tickets

arXiv:2607. 20555v1 Announce Type: new Abstract: The lottery ticket hypothesis proposes that large random neural networks contain sparse subnetworks that can match the performance of dense models after comparable training.

Hugging Face Trending Papers
Jun 10

Finding Sparse Subnetworks in One Training Cycle via Progressive Magnitude-Based Pruning

Neural network pruning reduces model size by removing less important parameters while aiming to preserve predictive performance. Although the Lottery Ticket Hypothesis (LTH) shows that sparse subnetworks can match dense networks when trained from suitable initializations, its iterative pruning procedure requires multiple complete training cycles.

arXiv Machine Learning
Aug 10

The Sparsity Whisperer

arXiv:2608. 06630v1 Announce Type: new Abstract: Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs.

By Linghao Kong, Inimai Subramanian, Micah Adler, Dan Alistarh, Dan Gutfreund, Nir Shavit
Hugging Face Trending Papers
Aug 19

A Unifying Relational Perspective on Expressive Lottery Tickets

The paper investigates how sparsity impacts the expressivity of graph neural networks, focusing on relational and temporal variants. It extends the Strong Expressive Lottery Ticket Hypothesis to multi-relational and temporal domains, proving that sufficiently large RGNNs contain sparse subnetworks that preserve 1‑RWL expressivity and providing a probabilistic bound for random pruning. Experiments validate the theoretical bounds, compare them to empirical results on synthetic data, and explore the relationship between pre‑training expressivity, optimization behavior, and prediction quality on temporal and molecular benchmarks.

arXiv AI
Jun 12

Structured vs. Unstructured Pruning: An Exponential Gap

arXiv:2603. 02234v3 Announce Type: replace-cross Abstract: The Strong Lottery Ticket Hypothesis (SLTH) states that large, randomly initialized neural networks contain sparse subnetworks capable of approximating a target function at initialization without training, suggesting that pruning alone is sufficient.

By Davide Ferre' (CNRS, COATI, UniCA, I3S), Fr\'ed\'eric Giroire (I3S, COATI, UniCA), Frederik Mallmann-Trenn (CNRS, COATI, I3S, UniCA), Emanuele Natale (CNRS, COATI, I3S, UniCA)
arXiv Machine Learning
Sep 21

Layerwise Decoupling for Stable Structured Sparsification of Fully Connected Layers

The paper introduces a layerwise, decoupled approach to structurally sparsify fully connected layers in pretrained neural networks. By extracting shallow two‑layer subnetworks, normalizing inner weights, and applying a structured group penalty to each block’s outer weight matrix, the method prunes neurons sequentially and reduces layer widths. The authors prove equivalence to a joint penalty for positively homogeneous activations, and demonstrate that this decoupled formulation is more robust, offering a broader regularization range and lower catastrophic over‑pruning while preserving accuracy in classification, sparse‑recovery, PINN, and OPT‑1.3B experiments.

By Charles Kulick, Armenak Petrosyan, Sui Tang