arXiv:2606. 07098v1 Announce Type: cross Abstract: We present SigmaScale, a method for learning auxiliary scaling matrices $S$ to aid truncated Singular Value Decomposition (SVD) based Large Language Model (LLM) compression.
By Ernests Lavrinovics, Marco Letizia, Roy Janco, Shai Segal, Johannes Bjerva, Maurizio Pierini
arXiv:2602.02848v2 Announce Type: replace
Abstract: Advances in large language models have driven strong performance across many tasks, but their memory and compute costs still hinder deployment. SVD...
By Ali Abbasi, Chayne Thrash, Haoran Qin, Shansita Sharma, Sepehr Seifi, Soheil Kolouri
The paper investigates how post‑training modifies the weights of Large Language Models relative to their pretrained state. By expressing weight updates in the pretrained matrix’s singular value decomposition, the authors separate changes into three geometric components: diagonal (reshaping singular values), off‑diagonal (rotating input‑output coupling), and null‑space (routing outside the original SVD core). Experiments on a math evaluation suite show that removing the diagonal component largely preserves post‑training gains, indicating that improvements stem mainly from reconfiguring and extending pretrained pathways rather than altering singular values.
By Jianing Qi, Hao Tang, Zhigang Zhu
arXiv:2602. 12429v2 Announce Type: replace Abstract: Foundation models have achieved remarkable success, yet their growing parameter counts pose significant computational and memory challenges.
By Paul Janson, Edouard Oyallon, Eugene Belilovsky
arXiv:2606. 06494v1 Announce Type: new Abstract: Parameter-efficient finetuning methods based on spectral decomposition have enabled progress in Continual Learning.
By Marius Dragoi, Ioana Pintilie, Alexandra Dragomir, Antonio Barbalau, Florin Brad
arXiv:2606. 19993v1 Announce Type: new Abstract: We present Activation- and Influence-Aware Ranks (AIR), an SVD-based LLM compression framework that guides each weight matrix's low-rank approximation with a backward-signal influence metric.
By Nico Harder, Daniel Becking, Karsten Mueller, Wojciech Samek
arXiv:2606.21847v2 Announce Type: replace-cross
Abstract: Low-rank decomposition is a promising compression paradigm for large language models (LLMs), yet its effectiveness hinges on rank budget allo...
By Chao Han, Yongjie Du, Junjie Tan, Zihao Xuan
arXiv:2608.23018v1 Announce Type: cross
Abstract: Federated fine-tuning of on-device large language models (LLMs) faces a significant computing burden. To overcome this limitation, split learning (SL...
By Tao Li, Yulin Tang, Qi Guo, Xianhao Chen
arXiv:2606. 16454v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) enables efficient adaptation of large pre-trained models to downstream tasks by parameterizing weight updates with low-rank matrices.
By Junghun Oh, Sungyong Baik, Kyoung Mu Lee
arXiv:2607. 03057v1 Announce Type: cross Abstract: The rapid growth in the parameter scale of large language models (LLMs) has created a strong demand for efficient compression techniques.
By Zhuowen Liu, Longkun Hao, Shiyu Feng, Xiaowen Chang, Ruiqun Li, Changqun Li
arXiv:2606. 31390v1 Announce Type: cross Abstract: Low-rank matrix optimization is often carried out via the Burer-Monteiro (BM) formulation, but choosing the factorization rank $r$ is delicate and can substantially slow optimization.
By Yudong Wei, Liang Zhang, Bingcong Li, Niao He
arXiv:2608. 01422v1 Announce Type: cross Abstract: Machine unlearning seeks to remove targeted information from trained models without requiring costly retraining.
By Tyler Lizzo, Larry Heck