arXiv Machine Learning By Ryan Solgi, Parsa Madinei, Jiayi Tian, Rupak Swaminathan, Jing Liu, Nathan Susanj, Zheng Zhang

Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM

Read the original on arXiv Machine Learning →

arXiv:2510. 05544v2 Announce Type: replace-cross Abstract: Large language models (LLM) and vision-language models (VLM) have achieved state-of-the-art performance, but they impose significant memory and computing challenges in deployment.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 2

LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models

arXiv:2606. 00573v1 Announce Type: new Abstract: Vision-language models (VLMs) deliver strong multimodal reasoning capabilities, but their large computational cost and high parameter counts make deployment challenging on resource-constrained devices.

By Haiyu Wang, Yutong Wang, Leshu Li, Yihui Ren, Sai Qian Zhang
arXiv AI
Sep 10

Mind the Approximation: Fisher-Weighted SVD Compression for ViTs

The paper introduces FACTS, a structured Fisher Approximation for compressing Vision Transformers (ViTs) using Fisher-weighted SVD, which enforces token‑local aggregation while preserving within‑token activation‑gradient dependence. It also presents Constrained Rank Search (CoRS) to optimize layer‑wise rank allocation under a fixed FLOP budget. Experiments on ViTs and hybrid architectures show that FACTS improves accuracy‑efficiency trade‑offs, outperforming the strongest SVD baseline by up to +5.8 percentage points on Swin‑B without requiring finetuning.

By Moritz Thoma, Maximilian Groezinger, Maximilian Forstenh\"ausler, Emad Aghajanzadeh, Ryan Pegoud, Manoj Rohit Vemparala, Pierpaolo Mori, Alexander Frickenstein, Daniel Mueller-Gritschneder, Ulf Schlichtmann