arXiv Machine Learning

Pruning Deep Neural Networks via the Marchenko--Pastur Distribution

arXiv:2606. 02608v1 Announce Type: new Abstract: We study a Marchenko--Pastur (MP) random-matrix approach to pruning deep neural networks with very small post-pruning fine-tuning budgets.

Hugging Face Trending Papers
Jun 10

Finding Sparse Subnetworks in One Training Cycle via Progressive Magnitude-Based Pruning

Neural network pruning reduces model size by removing less important parameters while aiming to preserve predictive performance. Although the Lottery Ticket Hypothesis (LTH) shows that sparse subnetworks can match dense networks when trained from suitable initializations, its iterative pruning procedure requires multiple complete training cycles.

arXiv Machine Learning
4d ago

Scaling Zero-Order Pretraining through Model Sharding

arXiv:2609.37899v1 Announce Type: new Abstract: Zero-order optimization (ZO) trains without backpropagation, making it relevant to forward-only hardware and non-differentiable loss, but its gradient...

By Francois Chaubard, Mykel J. Kochenderfer, Chris R\'e
arXiv Machine Learning
Sep 11

LILA: Calibration-Free Structured Pruning of Large Language Models via Latent Spectral Geometry

LILA (Latent-Informed Layer Analysis) introduces a calibration‑free method for structured pruning of large language models by scoring neuron importance using the Kolmogorov–Smirnov distance between singular value distributions of full and neuron‑ablated feed‑forward network weight matrices. The approach requires no training, calibration data, or auxiliary networks, and outperforms existing methods such as PruneNet and SliceGPT on LLaMA‑2‑7B and Phi‑2 at various sparsity levels. After a single epoch of LoRA fine‑tuning, LILA matches heavily calibrated baselines, and a Neural Tangent Kernel analysis provides theoretical support for its spectral importance criterion. Additionally, LILA can dynamically allocate sparsity budgets, achieving state‑of‑the‑art generative preservation and revealing architectural bottlenecks at higher compression.

By Sankar Behera, Dhruv Singh, Anshika Agnihotri, Raj Kumar Choudhary, Satyadev Ahlawat, Yamuna Prasad
arXiv Machine Learning
Sep 25

Spectral-Guided Diffusion: Accelerating Inference via Static Spectral Layer Scheduling

Spectral-Guided Diffusion introduces a method to accelerate diffusion inference by identifying and reusing residual branches that need not be recomputed during the trajectory. The approach uses a Spectral Concentration Ratio (SCR) combined with Frobenius magnitude to create an offline sensitivity proxy and deterministic lifetime for each scheduled unit, eliminating the need for routers or input-dependent searches. Experiments on models such as LLaDA-8B, DiT-XL/2, U-ViT-L, and SDXL show that this scheduling preserves quality better than several baselines and achieves up to a 3.0× wall‑clock speedup over eager inference.

By Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Anuj Sharma
arXiv Computer Vision
Sep 15

Sparsity-Adaptive Sharpness-Aware Minimization

The paper introduces Sparsity-Adaptive Sharpness-Aware Minimization (SA‑SAM), a method that adjusts the perturbation radius in sharpness-aware training to remain consistent as model sparsity increases. It also evaluates a Magnitude‑Weighted Hessian (MWH) importance metric derived from second‑order analysis. Experiments on CIFAR‑10‑C, CIFAR‑100‑C, and ImageNet‑100‑C show that SA‑SAM improves corruption robustness at 80–90% sparsity while maintaining clean accuracy, and the study reports inference throughput at deployment‑relevant sparsity levels.

By Shiryu Ueno, Yoshikazu Hayashi, Kunihito Kato
arXiv Machine Learning
Sep 25

Task-Aware Spectral Pruning: A Mixture-of-Masks Framework for Efficient LLM Inference

Task-Aware Spectral Pruning (TASP) is a post‑training framework that tailors sparse masks to specific tasks by calibrating module‑level spectral descriptors against task‑specific ablation effects. It constructs masks that close grouped‑query‑attention and SwiGLU dependencies, routing each user turn to a single compiled mask that remains fixed during prefill and decoding. In experiments, TASP achieves a 43% active‑FLOP reduction while preserving 97.7% of the dense BF16 performance on Llama‑3‑70B, and delivers a 1.44× speedup on an A100 80GB with INT8‑weight/BF16‑compute, reducing decode latency from 45.2 to 31.3 ms/token.

By Ibne Farabi Shihab, Fariya Afrin, Sanjeda Akter, Anuj Sharma