The paper introduces Global Relative Kinetic Utility (Global RKU), a label‑free method for calibrating cross‑layer credit in global structured pruning of large language models. Global RKU estimates channel importance via a final‑hidden‑state activation‑gradient signal and applies block‑relative normalization to remove block‑common scale while preserving within‑block ordering, enabling a single‑stage static pruning topology. Experiments on Qwen‑2.5‑7B show significant performance gains at various sparsity levels, and ablation studies confirm the effectiveness of the relative‑normalization step.
By Tianhao Qian, Guilin Qi, Jiayu Chen
Neural network pruning reduces model size by removing less important parameters while aiming to preserve predictive performance. Although the Lottery Ticket Hypothesis (LTH) shows that sparse subnetworks can match dense networks when trained from suitable initializations, its iterative pruning procedure requires multiple complete training cycles.
arXiv:2606. 12278v1 Announce Type: cross Abstract: Neural network pruning reduces model size by removing less important parameters while aiming to preserve predictive performance.
By Romana Qureshi, Hafida Benhidour, Said Kerrache, Nahlah Aljeraisy
The paper introduces a Hybrid Quadratic Unconstrained Binary Optimization (QUBO) framework for structured neural network pruning that integrates task‑aware sensitivity metrics (first‑order Taylor and Weight‑Fisher) into the objective’s linear term and optionally uses activation similarity for quadratic interactions. It controls pruning cardinality via a binary search over a capacity incentive rather than an explicit penalty and further refines the pruning mask with a two‑stage QUBO–Tensor‑Train strategy that employs gradient‑free black‑box optimization. Experiments on SIDD image denoising with a Half‑UNet model demonstrate that this Hybrid QUBO outperforms Taylor and L1‑based QUBO baselines in PSNR and SSIM, while also revealing computational and deployment challenges of mask‑based pruning.
By Osama Orabi, Artur Zagitov, Hadi Salloum, Viktor A. Lobachev, Yaroslav Kholodov
arXiv:2602. 24266v2 Announce Type: replace-cross Abstract: Which internal mechanisms of a neural network can be replaced while preserving the computation it performs?
By Amir Asiaee
arXiv:2608. 10989v1 Announce Type: cross Abstract: Token-pruning policies are usually designed for a single recognition pipeline, but pretrained Vision Transformers are reused across tasks with different spatial demands.
By Hongsen Cao, Mona Jaber, Shanxin Yuan, Ahmed Sayed
COEC (Calibrated Orthogonal-Equivalence Compensation) is a training‑free framework that improves structured pruning of large language models by applying alternating left and right orthogonal rotations to the retained weight matrix. The method optimizes the right rotation on a reduced Stiefel manifold, rescales singular values via generalized cross‑validation, tempers the calibration Gram matrix, and adds an alignment penalty to preserve geometric relations between attention projections. Experiments on Llama‑3, Llama‑3.1, and Qwen2.5 show that COEC consistently improves perplexity and zero‑shot accuracy across multiple sparsity levels, outperforming existing compensation techniques.
By Peiqi Yu, Nam Ling, Wei Wang, Wei Jiang
arXiv:2608.23253v1 Announce Type: cross
Abstract: Vision-language models typically encode an image into hundreds of visual tokens, incurring substantial inference latency and GPU memory overhead. Exi...
By Taoyu Qian, Qi Wang, Daqian Shi, Yuanhao Jiang, Shang Gao, Hualong Yu
arXiv:2608. 06901v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved remarkable generalization across diverse multimodal tasks through large-scale pre-training, yet their rapidly increasing computational and memory requirements pose significant challenges for deployment in constrained environments.
By Minseok Kang, Hyunwoo Kim, Chanyoung Kim, Minwoo Kim, Jaekoo Lee, Dahuin Jung
arXiv:2603. 12222v2 Announce Type: replace-cross Abstract: Vision Transformers require significant computational resources and memory bandwidth, severely limiting their deployment on resource-constraint hardware.
By Andy Li, Aiden Durrant, Milan Markovic, Georgios Leontidis
Approximate machine unlearning seeks to remove the influence of a forget set from a trained model without full retraining. Existing gradient-based methods require data-dependent hyperparameter search,...
The paper introduces Unmerge, an efficient machine unlearning algorithm that treats unlearning as the inverse of task arithmetic. By representing the forget component as a low‑rank basis at each layer, Unmerge optimizes three goals—matching the merged vector, suppressing leakage, and bounding correction size—to limit forget leakage and retain damage. Experiments on ResNet‑50, ViT‑S/16, and Llama‑3.2‑3B show significant performance gains over existing methods while maintaining privacy and feature‑distribution fidelity.
By Haoran Tang, Andrew Tan, Rajiv Khanna