arXiv:2606. 10445v1 Announce Type: new Abstract: Semi-structured 2:4 sparsity is widely supported by modern accelerators, providing up to a 2x theoretical speedup.
By Jaeseong Lee, Seung-won Hwang, Samyam Rajbhandari
Semi-structured 2:4 sparsity is widely supported by modern accelerators, providing up to a 2x theoretical speedup. However, its strict 50% sparsity constraint often causes non-negligible accuracy degradation under post-training pruning.
arXiv:2607. 28418v1 Announce Type: cross Abstract: Pruning is a promising approach for improving the efficiency of LLMs.
By Haozhe Hu, Hao Wu, Peiran Yin, Chao Han, Yunpu Ma, Xiaoyu Shen
arXiv:2605. 17289v2 Announce Type: replace-cross Abstract: Unstructured sparsity is now natively accelerated by recent GPU kernels and dataflow hardware, shifting the bottleneck from inference execution to the pruning algorithm.
By Mohammad Mozaffari, Younes Hourri, Mohammad Rastegari, Mahyar Najibi
arXiv:2609.10311v1 Announce Type: cross
Abstract: The lottery ticket hypothesis posits the existence of winning tickets: sparse subnetworks that, when trained in isolation from their original initial...
By Benedikt Tscheschner, Eduardo Veas, Marc Masana
The paper introduces Soft-OMP and Soft-IHT, permutation‑based variants of Orthogonal Matching Pursuit and Iterative Hard Thresholding that replace the non‑differentiable argsort with continuous soft‑sort operators. These differentiable algorithms enable the construction of fully trainable neural network architectures—OMP‑Net and IHT‑Net—while preserving the core greedy sparse recovery logic. The authors show both theoretically and numerically that the soft variants approximate their hard counterparts with controllable accuracy and can be extended to structured sparse recovery by learning structure‑aware weights.
By Sina Mohammad-Taheri, Matthew J. Colbrook, Simone Brugiapaglia