arXiv Machine Learning By Ivan Ilin, Peter Richt\'arik

QK-Wanda: Coupling Queries and Keys for Unstructured Pruning

Read the original on arXiv Machine Learning →

QK-Wanda is a new pruning method that extends Wanda by scoring query and key weights together, using each projection’s deletion cost under an unmasked pre‑RoPE reconstruction objective. It augments each projection’s score with information from the opposite projection, enabling shared pruning budgets without requiring gradients or weight updates. Across 15 models ranging from 0.5 B to 72 B parameters, QK-Wanda reduces reconstruction error by 60 % at 50 % sparsity and 45 % at 80 % compared to Wanda, and yields notable downstream improvements on some models such as Llama 2 70 B. whyItMatters":"The study demonstrates that coupling query and key pruning criteria can substantially lower reconstruction error and improve downstream performance, highlighting both the potential and limitations of local reconstruction as a predictor of model quality."

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 24

COEC: Calibrated Orthogonal-Equivalence Compensation for Structured Pruning of Large Language Models

COEC (Calibrated Orthogonal-Equivalence Compensation) is a training‑free framework that improves structured pruning of large language models by applying alternating left and right orthogonal rotations to the retained weight matrix. The method optimizes the right rotation on a reduced Stiefel manifold, rescales singular values via generalized cross‑validation, tempers the calibration Gram matrix, and adds an alignment penalty to preserve geometric relations between attention projections. Experiments on Llama‑3, Llama‑3.1, and Qwen2.5 show that COEC consistently improves perplexity and zero‑shot accuracy across multiple sparsity levels, outperforming existing compensation techniques.

By Peiqi Yu, Nam Ling, Wei Wang, Wei Jiang
arXiv Machine Learning
Sep 25

Task-Aware Spectral Pruning: A Mixture-of-Masks Framework for Efficient LLM Inference

Task-Aware Spectral Pruning (TASP) is a post‑training framework that tailors sparse masks to specific tasks by calibrating module‑level spectral descriptors against task‑specific ablation effects. It constructs masks that close grouped‑query‑attention and SwiGLU dependencies, routing each user turn to a single compiled mask that remains fixed during prefill and decoding. In experiments, TASP achieves a 43% active‑FLOP reduction while preserving 97.7% of the dense BF16 performance on Llama‑3‑70B, and delivers a 1.44× speedup on an A100 80GB with INT8‑weight/BF16‑compute, reducing decode latency from 45.2 to 31.3 ms/token.

By Ibne Farabi Shihab, Fariya Afrin, Sanjeda Akter, Anuj Sharma