Unified Response Geometry for Structured Pruning
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The paper introduces Global Relative Kinetic Utility (Global RKU), a label‑free method for calibrating cross‑layer credit in global structured pruning of large language models. Global RKU estimates channel importance via a final‑hidden‑state activation‑gradient signal and applies block‑relative normalization to remove block‑common scale while preserving within‑block ordering, enabling a single‑stage static pruning topology. Experiments on Qwen‑2.5‑7B show significant performance gains at various sparsity levels, and ablation studies confirm the effectiveness of the relative‑normalization step.
Neural network pruning reduces model size by removing less important parameters while aiming to preserve predictive performance. Although the Lottery Ticket Hypothesis (LTH) shows that sparse subnetworks can match dense networks when trained from suitable initializations, its iterative pruning procedure requires multiple complete training cycles.
arXiv:2606. 12278v1 Announce Type: cross Abstract: Neural network pruning reduces model size by removing less important parameters while aiming to preserve predictive performance.
The paper introduces a Hybrid Quadratic Unconstrained Binary Optimization (QUBO) framework for structured neural network pruning that integrates task‑aware sensitivity metrics (first‑order Taylor and Weight‑Fisher) into the objective’s linear term and optionally uses activation similarity for quadratic interactions. It controls pruning cardinality via a binary search over a capacity incentive rather than an explicit penalty and further refines the pruning mask with a two‑stage QUBO–Tensor‑Train strategy that employs gradient‑free black‑box optimization. Experiments on SIDD image denoising with a Half‑UNet model demonstrate that this Hybrid QUBO outperforms Taylor and L1‑based QUBO baselines in PSNR and SSIM, while also revealing computational and deployment challenges of mask‑based pruning.
arXiv:2602. 24266v2 Announce Type: replace-cross Abstract: Which internal mechanisms of a neural network can be replaced while preserving the computation it performs?
arXiv:2608. 10989v1 Announce Type: cross Abstract: Token-pruning policies are usually designed for a single recognition pipeline, but pretrained Vision Transformers are reused across tasks with different spatial demands.