arXiv Machine Learning

Deep greedy unfolding: Sorting out argsorting in greedy sparse recovery algorithms

The paper introduces Soft-OMP and Soft-IHT, permutation‑based variants of Orthogonal Matching Pursuit and Iterative Hard Thresholding that replace the non‑differentiable argsort with continuous soft‑sort operators. These differentiable algorithms enable the construction of fully trainable neural network architectures—OMP‑Net and IHT‑Net—while preserving the core greedy sparse recovery logic. The authors show both theoretically and numerically that the soft variants approximate their hard counterparts with controllable accuracy and can be extended to structured sparse recovery by learning structure‑aware weights.

arXiv Machine Learning
4d ago

OMP-MoE: Efficient Expert Pruning for Mixture-of-Experts LLMs via Orthogonal Matching Pursuit

OMP-MoE is a training‑free compression framework that prunes redundant experts in Mixture‑of‑Experts large language models by framing the problem as sparse signal reconstruction solved with Orthogonal Matching Pursuit. The method greedily selects expert contributions as dictionary atoms to minimize reconstruction error, then optimizes cross‑layer expert allocation via a water‑filling strategy, and finally introduces an adaptive inference mechanism (OMP‑MoE†) that dynamically adjusts expert activation based on energy prediction. Experiments on Qwen, DeepSeek‑V2, GPT‑OSS, and Mixtral MoE show consistent performance gains at 25‑50% pruning ratios, with Qwen3‑30B‑A3B retaining 93.3% of original performance at 50% compression while achieving significant speedups.

By Dezhi Li, Lujun Li, Qiyuan Zhu, Hao Gu, Bei Liu, Sirui Han, Yike Guo
arXiv Machine Learning
Sep 21

Layerwise Decoupling for Stable Structured Sparsification of Fully Connected Layers

The paper introduces a layerwise, decoupled approach to structurally sparsify fully connected layers in pretrained neural networks. By extracting shallow two‑layer subnetworks, normalizing inner weights, and applying a structured group penalty to each block’s outer weight matrix, the method prunes neurons sequentially and reduces layer widths. The authors prove equivalence to a joint penalty for positively homogeneous activations, and demonstrate that this decoupled formulation is more robust, offering a broader regularization range and lower catastrophic over‑pruning while preserving accuracy in classification, sparse‑recovery, PINN, and OPT‑1.3B experiments.

By Charles Kulick, Armenak Petrosyan, Sui Tang
arXiv Machine Learning
Jun 2

Learning-Augmented Scalable Linear Assignment Problem Optimization via Neural Dual Warm-Starts

arXiv:2605. 09382v2 Announce Type: replace Abstract: The Linear Assignment Problem is a fundamental combinatorial optimization task where classical exact solvers ensure optimality but suffer from an $\mathcal{O}(N^{3})$ bottleneck, while recent neural approximations struggle with scalability and exactness.

By Ilay Yavlovich, Jad Agbaria, Muhamed Mhamed, Nir Weinberger, Jose Yallouz
arXiv AI
Aug 18

CG-GLORE: A Conjugate Gradient-Based Global-Local Regularization Network for Sparse-View CT Reconstruction

arXiv:2608. 15246v1 Announce Type: cross Abstract: Sparse-view computed tomography (CT) reduces radiation dose by acquiring fewer projection views, but the resulting inverse problem is highly ill-posed and often produces severe streak artifacts.

By Tran Xuan Hieu Le, Doanh C. Bui, Vu Trung Duong Le, Hoai Luan Pham, Khang Nguyen, Mai K. Nguyen, Tu Bao Ho, Yasuhiko Nakashima
arXiv Machine Learning
Aug 10

The Sparsity Whisperer

arXiv:2608. 06630v1 Announce Type: new Abstract: Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs.

By Linghao Kong, Inimai Subramanian, Micah Adler, Dan Alistarh, Dan Gutfreund, Nir Shavit
arXiv Computer Vision
Sep 15

Newton Deep Unfolding for Compressed Sensing

arXiv:2609.14391v1 Announce Type: new Abstract: Compressed sensing (CS) reconstructs images from highly limited measurements, but existing deep unfolding methods are typically driven by first-order o...

By Changhua He, Xianchao Xiu
arXiv AI
Jun 9

Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs

arXiv:2605. 15491v2 Announce Type: replace-cross Abstract: Layer pruning removes entire Transformer decoder blocks from large language models, but introduces a mismatch between the hidden state received by the next surviving layer and the distribution it was trained to process, leading to significant performance degradation.

By Vincent-Daniel Yun, Junhyuk Jo, Sai Praneeth Karimireddy, Sunwoo Lee