arXiv:2607. 03784v1 Announce Type: cross Abstract: While prior studies have successfully compressed vision Transformers (ViTs) through various pruning techniques, most have concentrated on width pruning to achieve significant reductions in model size.
By Zhenfeng Su, Kang Zhao, Han Bao, Tao Yuan, Zhongzhe Hu, Xianzhi Yu, Wenxuan Wang
arXiv:2603. 12222v2 Announce Type: replace-cross Abstract: Vision Transformers require significant computational resources and memory bandwidth, severely limiting their deployment on resource-constraint hardware.
By Andy Li, Aiden Durrant, Milan Markovic, Georgios Leontidis
arXiv:2606. 03428v1 Announce Type: cross Abstract: The large sizes of Spiking Vision Transformers (SViTs) still hinder their embedded implementation, highlighting the need for model compression.
By Rachmad Vidya Wicaksana Putra, Achyuta Muthuvelan, Alberto Marchisio, Muhammad Shafique
arXiv:2606. 03257v1 Announce Type: cross Abstract: Spiking Vision Transformer (SViT) models are promising low-power ViT models for solving vision-based tasks with state-of-the-art performance.
By Rachmad Vidya Wicaksana Putra, Achyuta Muthuvelan, Alberto Marchisio, Muhammad Shafique
The paper introduces a hardware‑aware framework that uses genetic programming to evolve layer‑specific scalar functions for Vision Transformers, replacing traditional LayerNorm with efficient, heterogeneous approximations. By applying a post‑training re‑alignment strategy, the method eliminates the need for full model retraining while achieving 90‑93% variance capture and recovering over 84% of ImageNet‑1K Top‑1 accuracy for ViT‑B and ViT‑L. The resulting architecture removes the global reduction bottleneck, reducing arithmetic complexity and off‑chip memory traffic, thereby enabling efficient deployment of ViTs on edge accelerators.
By Kieran Carrigg, Sigur de Vries, Amirhossein Sadough, Marcel van Gerven
arXiv:2608. 06901v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved remarkable generalization across diverse multimodal tasks through large-scale pre-training, yet their rapidly increasing computational and memory requirements pose significant challenges for deployment in constrained environments.
By Minseok Kang, Hyunwoo Kim, Chanyoung Kim, Minwoo Kim, Jaekoo Lee, Dahuin Jung