Magnitude Profile Pruning introduces a training‑free, calibration‑free method for removing attention heads in Transformer models by statistically detecting outliers in weight row norms. Heads whose projection weights fall within the bulk of the distribution are pruned, while outlier heads are retained. Across several models, the MP‑G variant achieves superior perplexity at various sparsity levels and yields significant parameter and FLOP reductions without requiring forward passes, calibration data, or gradient computations.
By Kasun Dewage, Marianna Pensky, Heranga K. Rathnasekara, Suranadi De Silva
arXiv:2607. 22587v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong performance across diverse tasks but their deployment is constrained by the memory and compute cost of their parameters.
By Manel Kara laoua, Soumia Bouyahiaoui, Aicha Boutorh
arXiv:2601.06787v2 Announce Type: replace
Abstract: Large Language Models (LLMs) are known to contain significant redundancy, yet a systematic explanation for why certain components, particularly in...
By Jaewon Sok, Jewon Yeom, Seonghyeon Park, Jeongjae Park, Taesup Kim
arXiv:2607. 22583v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved widespread adoption because of their strong reasoning and query-response capabilities.
By Muhammad Junaid Ali, Smail Niar, El-Ghazali Talbi
arXiv:2606. 02559v1 Announce Type: cross Abstract: Post-training compression of Large Language Models (LLMs) removes entire architectural components, either deleting them or replacing them with fitted modules.
By Elia Cunegatti, Marcus Vukojevic, Erik Nielsen, Giovanni Iacca
arXiv:2606. 18304v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models scale compute efficiently, yet remain expensive to deploy due to their substantial memory footprint and inference overhead.
By Yifu Ding, Jiacheng Wang, Ge Yang, Yongcheng Jing, Jinyang Guo, Xianglong Liu, Dacheng Tao
arXiv:2606. 24173v1 Announce Type: cross Abstract: On-device fault detection enables real-time diagnostics without cloud dependency, but deploying machine learning models on resource-constrained hardware demands careful tradeoffs between accuracy, latency, and model size.
By Disha Patel
arXiv:2605. 18331v2 Announce Type: replace Abstract: Large Language Models (LLMs) have experienced significant growth and development in recent years.
By Diego Coello de Portugal Mecke, Tom Hanika, Lars Schmidth-Thieme
arXiv:2606. 03428v1 Announce Type: cross Abstract: The large sizes of Spiking Vision Transformers (SViTs) still hinder their embedded implementation, highlighting the need for model compression.
By Rachmad Vidya Wicaksana Putra, Achyuta Muthuvelan, Alberto Marchisio, Muhammad Shafique
Task-Aware Spectral Pruning (TASP) is a post‑training framework that tailors sparse masks to specific tasks by calibrating module‑level spectral descriptors against task‑specific ablation effects. It constructs masks that close grouped‑query‑attention and SwiGLU dependencies, routing each user turn to a single compiled mask that remains fixed during prefill and decoding. In experiments, TASP achieves a 43% active‑FLOP reduction while preserving 97.7% of the dense BF16 performance on Llama‑3‑70B, and delivers a 1.44× speedup on an A100 80GB with INT8‑weight/BF16‑compute, reducing decode latency from 45.2 to 31.3 ms/token.
By Ibne Farabi Shihab, Fariya Afrin, Sanjeda Akter, Anuj Sharma
arXiv:2512.20636v2 Announce Type: replace-cross
Abstract: Many self-attention sublayers in large language models (LLMs) can be removed with little to no loss. We attribute this to the Attention Suppr...
By Dhananjay Saikumar, Blesson Varghese
arXiv:2608. 05499v1 Announce Type: cross Abstract: Modern deep neural networks achieve strong performance, but their scale makes them costly and slow, especially on resource-constrained edge devices.
By Sadegh Jafari, Mohiuddin Bilwal, Fan Zhou, Brian Gelder, Ali Jannesari