arXiv:2606. 00573v1 Announce Type: new Abstract: Vision-language models (VLMs) deliver strong multimodal reasoning capabilities, but their large computational cost and high parameter counts make deployment challenging on resource-constrained devices.
By Haiyu Wang, Yutong Wang, Leshu Li, Yihui Ren, Sai Qian Zhang
arXiv:2411. 09816v5 Announce Type: replace Abstract: Large neural networks achieve state-of-the-art performance on many tasks, yet their sheer size hinders deployment on resource-constrained devices.
By Cem \"Uy\"uk, Mike Lasby, Mohamed Yassin, Utku Evci, Yani Ioannou
The paper introduces FACTS, a structured Fisher Approximation for compressing Vision Transformers (ViTs) using Fisher-weighted SVD, which enforces token‑local aggregation while preserving within‑token activation‑gradient dependence. It also presents Constrained Rank Search (CoRS) to optimize layer‑wise rank allocation under a fixed FLOP budget. Experiments on ViTs and hybrid architectures show that FACTS improves accuracy‑efficiency trade‑offs, outperforming the strongest SVD baseline by up to +5.8 percentage points on Swin‑B without requiring finetuning.
By Moritz Thoma, Maximilian Groezinger, Maximilian Forstenh\"ausler, Emad Aghajanzadeh, Ryan Pegoud, Manoj Rohit Vemparala, Pierpaolo Mori, Alexander Frickenstein, Daniel Mueller-Gritschneder, Ulf Schlichtmann
arXiv:2510. 05544v2 Announce Type: replace-cross Abstract: Large language models (LLM) and vision-language models (VLM) have achieved state-of-the-art performance, but they impose significant memory and computing challenges in deployment.
By Ryan Solgi, Parsa Madinei, Jiayi Tian, Rupak Swaminathan, Jing Liu, Nathan Susanj, Zheng Zhang
The paper introduces FrameFT, a parameter-efficient fine-tuning method for transformer models that represents weight updates using sparse coefficients in a Fusion Frame basis. This approach reduces memory usage by storing only the sparse coefficients, leading to significant compute advantages and formal convergence guarantees. Experiments on language and vision tasks show that FrameFT matches or surpasses state‑of‑the‑art PEFT techniques while requiring far fewer trainable parameters.
By Harshavardhan Adepu, Li Zhang, Sanjiv Kumar, Vikas Singh
The paper introduces Channel Group-Shared (CGS) low‑rank approximation, a Singular Value Decomposition–based strategy that shares down/up‑projection matrices across channel groups while using lightweight diagonal matrices for each group. This design dramatically cuts the parameter count of pointwise convolutions, which dominate the size of large‑kernel CNNs such as RepLKNet, ConvNeXt, and SLaK. Experiments show that CGS‑enhanced models maintain competitive accuracy while substantially reducing storage, memory bandwidth, and loading latency, making them viable for deployment on edge devices.
By Hao Luo, Yiting Yang, Wenyi Zhao, Man Jiang, Zhijun Lin, Ghulam Mohiuddin, Ting Jiang, Kunming Luo, Zihao Zhang, Qingsen Yan, Guoqing Wang, Wei Dong, Peng Wang