arXiv Machine Learning

Train Large, Deploy Compact: Structured Compression for Compact Low-Rank Adaptation

arXiv:2510. 00192v3 Announce Type: replace Abstract: Low-rank adaptation (LoRA) has become a widely used paradigm for parameter-efficient fine-tuning of large language models, yet its representational capacity often lags behind full fine-tuning.

arXiv AI
Jul 21

Sparsity-Aware Low-Rank Representation for Efficient Fine-Tuning of Large Language Models

arXiv:2601. 16991v3 Announce Type: replace-cross Abstract: Adapting large pre-trained language models to downstream tasks often entails fine-tuning millions of parameters or deploying costly dense weight updates, which hinders their use in resource-constrained environments.

By Longteng Zhang, Sen Wu, Shuai Hou, Zhengyu Qing, Zhuo Zheng, Danning Ke, Qihong Lin, Qiang Wang, Shaohuai Shi, Xiaowen Chu
Hugging Face Trending Papers
Sep 24

Automatic Rank Allocation for Low-Rank Adaptation in Large Language Models via lp Regularization

The paper introduces αp-LoRA, a rank-allocation strategy for low-rank adaptation (LoRA) in large language models that uses π-regularization (0 < p < 1) to induce sparsity in rank-one components. By regularizing the energy of each component, redundant parts are encouraged to vanish while important ones are retained, and the authors derive a proximal subproblem that reduces the matrix optimization to a two‑dimensional thresholding criterion. Experiments on natural language understanding and question‑answering tasks show that αp-LoRA achieves performance competitive with existing LoRA baselines.

arXiv Computation and Language
Aug 31

Pruning Laws for Large Language Models

arXiv:2504.04342v2 Announce Type: replace Abstract: Scaling up model parameters and training data consistently improves the performance of large language models (LLMs), but at the cost of rapidly gro...

By Ayan Sengupta, Siddhant Chaudhary, Tanmoy Chakraborty
arXiv Machine Learning
Sep 25

Automatic Rank Allocation for Low-Rank Adaptation in Large Language Models via lp Regularization

The paper introduces ρ_p-LoRA, a rank-allocation strategy for low-rank adaptation (LoRA) in large language models that uses ρ_p regularization (0 < p < 1) to induce sparsity in rank-one components. By regularizing the energy of each component, redundant parts are encouraged to vanish while important ones are retained, leading to an implicit thresholding criterion derived from a two-dimensional proximal subproblem. Experiments on natural language understanding and question-answering tasks show that ρ_p-LoRA achieves performance comparable to existing LoRA baselines.

By Zebang Xie, Chuanyang Zheng, Yik-Chung Wu, Yihang Gao
arXiv Machine Learning
Jul 21

NIRVANA: Structured Pruning Reimagined for Large Language Model Compression

arXiv:2509. 14230v2 Announce Type: replace Abstract: While structured pruning presents a highly effective pathway for accelerating Large Language Model (LLM) inference, existing methods frequently suffer from significant performance degradation and demand computationally retraining to recover capabilities.

By Mengting Ai, Tianxin Wei, Sirui Chen, Jingrui He
arXiv Machine Learning
Aug 27

Resource-Efficient Pruning for Transformer via Low-Rank Importance Estimation

The paper introduces REP‑LIE, a resource‑efficient pruning method for Transformer models that estimates weight importance using gradients from LoRA low‑rank matrices, avoiding full gradient computation. It incorporates a stability score to iteratively prune unimportant parameters and then fine‑tunes the pruned model with lightweight updates, eliminating the need for full‑parameter optimization. Experiments on medium‑scale encoders and large‑scale generative models such as LLaMA‑7B and Mistral‑7B show that REP‑LIE achieves competitive performance compared to existing pruning approaches.

By Peng Liu, Huibing Zeng, Yiqun Zhang, Yang Yi, Jigang Wu