arXiv Machine Learning

Can Spectral-Clipping Enable Better Learning While Forgetting Less for Low-Rank Adaptation?

arXiv:2608. 12332v1 Announce Type: cross Abstract: In recent years, low-rank adaptation (LoRA) has emerged as a significant paradigm that freezes pre-trained weights and introduces small, learnable adapters instead of fine-tuning the full set of parameters.

arXiv AI
3d ago

Foundation-Preserving Optimization in Generalized Eigenspace

The paper introduces Foundation Preserving LoRA (FoLoRA), a forgetting‑aware optimization framework that balances adaptation to downstream tasks with preservation of pretraining capabilities. FoLoRA uses a first‑order preservation condition to define a forgetting penalty based on pretraining‑proxy activations and a task utility from downstream activations, scoring update directions via a generalized Rayleigh quotient. This spectral coordinate system enables gated Adam updates that reduce low‑utility, high‑penalty directions, and the method constructs pretraining proxy calibration data by sampling from the pretrained model. Experiments on math, code, and instruction‑following tasks demonstrate that FoLoRA achieves a stronger balance between target task performance and aggregate preservation of non‑target capabilities compared to baselines.

By Dongjun Kim, Adrian de Wynter, Huancheng Chen, Heasung Kim, Haris Vikalo
arXiv Machine Learning
1d ago

Learn the Directions, Normalize the Gains: Post-Training Normalization for LoRA

The paper introduces LoRA‑Norm, a post‑training normalization technique for Low‑Rank Adaptation (LoRA) that rebalances the gains of learned singular directions without altering the directions themselves. LoRA‑Norm uses spectral rebalancing and nuclear‑norm restoration to preserve total spectral mass, requiring no calibration data or extra training and adding no inference overhead. Experiments on two backbones and three adaptation tasks show that LoRA‑Norm improves both specialization and capability retention, outperforming other post‑hoc spectral pruning and gradient‑guided editing methods.

By Zailong Tian, Yanzhe Chen, Zhuoheng Han, Houfeng Wang, Lizi Liao
arXiv AI
2d ago

Local Support Learning

arXiv:2610.02126v1 Announce Type: cross Abstract: We explore catastrophic forgetting in the context of large pre-trained models. By considering forgetting as a geometric problem in the input space of...

By Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes
arXiv Machine Learning
Sep 1

Normalized Low-Rank Adaptation

Normalized Low-Rank Adaptation (NoRA) is a lightweight enhancement to the widely used LoRA technique that normalizes the down‑projection matrices during training. By doing so, NoRA stabilizes early optimization dynamics, accelerates convergence, and improves performance across pretraining, supervised fine‑tuning, and reinforcement learning. The method adds no extra trainable parameters or inference‑time cost, making it broadly applicable.

By Jiale Kang, Ziyin Yue, Zheng Zhan, Yangyi Huang, Weiyang Liu