arXiv AI By Viktoriia Chekalina, Daniil Moskovskiy, Tatiana Matveeva, Andrey Kuznetsov, Evgeny Frolov

Generalized Fisher-Weighted SVD: Scalable Kronecker-Factored Fisher Approximation for Compressing Large Language Models

Read the original on arXiv AI →

arXiv:2505. 17974v2 Announce Type: replace-cross Abstract: The Fisher information is a fundamental concept for characterizing the sensitivity of parameters in neural networks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 10

Mind the Approximation: Fisher-Weighted SVD Compression for ViTs

The paper introduces FACTS, a structured Fisher Approximation for compressing Vision Transformers (ViTs) using Fisher-weighted SVD, which enforces token‑local aggregation while preserving within‑token activation‑gradient dependence. It also presents Constrained Rank Search (CoRS) to optimize layer‑wise rank allocation under a fixed FLOP budget. Experiments on ViTs and hybrid architectures show that FACTS improves accuracy‑efficiency trade‑offs, outperforming the strongest SVD baseline by up to +5.8 percentage points on Swin‑B without requiring finetuning.

By Moritz Thoma, Maximilian Groezinger, Maximilian Forstenh\"ausler, Emad Aghajanzadeh, Ryan Pegoud, Manoj Rohit Vemparala, Pierpaolo Mori, Alexander Frickenstein, Daniel Mueller-Gritschneder, Ulf Schlichtmann
arXiv AI
Sep 3

Scalable Kronecker-Fisher Approximation: Efficient Hessian Analysis for Billion-Parameter Language Models Compression

The paper introduces a scalable Kronecker-based approximation that captures cross-layer interactions without storing the full Fisher matrix, making Hessian analysis feasible for billion-parameter language models. It identifies consistent vulnerability patterns, notably that value projection layers are the most sensitive and exhibit strong cross-layer correlations across various model families. Experiments on quantization, sparsification, inter-layer corruption, and fine-tuning show that the approximation correlates strongly with performance degradation and recovery, providing a practical tool for identifying fragile components and guiding compression and optimization strategies.

By Viacheslav Yusupov, Daria Cherniuk, Evgeny Frolov