Structured Neuron Pruning in Deep Neural Networks Using Multi-Armed Bandits
arXiv:2606. 07615v1 Announce Type: cross Abstract: Deep neural networks often contain redundant hidden units.
arXiv:2607. 22564v1 Announce Type: new Abstract: Convolutional neural networks often contain redundant feature maps that increase storage and inference cost.
arXiv:2606. 07615v1 Announce Type: cross Abstract: Deep neural networks often contain redundant hidden units.
The paper introduces a structured post‑training pruning method for vision and language transformers called Damage‑Aware Bandit Pruning. It treats the selection of functional units (attention heads and MLP channel groups) as a multi‑armed bandit problem, using paired damage (masked loss minus base loss) as a reward to guide either UCB or Thompson Sampling policies. Experiments on a range of models (GPT‑2, OPT, Pythia, Qwen2.5, SmolLM2, ViT‑B/16, DeiT‑Tiny, Swin‑Tiny) show that the bandit approaches generally reduce degradation compared to budgeted‑greedy baselines, with statistically significant improvements in most comparisons.
arXiv:2606. 08574v1 Announce Type: new Abstract: Data pruning (DP), as an oft-stated strategy to alleviate heavy training burdens, reduces the volume of training samples according to a well-defined pruning method while striving for near-lossless performance.
arXiv:2609.10346v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) process hundreds or thousands of visual tokens per image, incurring prohibitive inference costs. While existin...
arXiv:2609.10311v1 Announce Type: cross Abstract: The lottery ticket hypothesis posits the existence of winning tickets: sparse subnetworks that, when trained in isolation from their original initial...
arXiv:2606. 12278v1 Announce Type: cross Abstract: Neural network pruning reduces model size by removing less important parameters while aiming to preserve predictive performance.
Neural network pruning reduces model size by removing less important parameters while aiming to preserve predictive performance. Although the Lottery Ticket Hypothesis (LTH) shows that sparse subnetworks can match dense networks when trained from suitable initializations, its iterative pruning procedure requires multiple complete training cycles.
The paper introduces a Hybrid Quadratic Unconstrained Binary Optimization (QUBO) framework for structured neural network pruning that integrates task‑aware sensitivity metrics (first‑order Taylor and Weight‑Fisher) into the objective’s linear term and optionally uses activation similarity for quadratic interactions. It controls pruning cardinality via a binary search over a capacity incentive rather than an explicit penalty and further refines the pruning mask with a two‑stage QUBO–Tensor‑Train strategy that employs gradient‑free black‑box optimization. Experiments on SIDD image denoising with a Half‑UNet model demonstrate that this Hybrid QUBO outperforms Taylor and L1‑based QUBO baselines in PSNR and SSIM, while also revealing computational and deployment challenges of mask‑based pruning.
arXiv:2606. 11761v1 Announce Type: new Abstract: Dynamic data pruning techniques aim to reduce computational cost while minimizing information loss by periodically selecting representative subsets of input data during model training.
arXiv:2604.13287v2 Announce Type: replace Abstract: Weight pruning is a common technique for compressing large neural networks. We focus on the challenging post-training one-shot setting, where a pre...
arXiv:2509. 14230v2 Announce Type: replace Abstract: While structured pruning presents a highly effective pathway for accelerating Large Language Model (LLM) inference, existing methods frequently suffer from significant performance degradation and demand computationally retraining to recover capabilities.
arXiv:2504.04342v2 Announce Type: replace Abstract: Scaling up model parameters and training data consistently improves the performance of large language models (LLMs), but at the cost of rapidly gro...