arXiv Machine Learning By Zailong Tian, Yanzhe Chen, Zhuoheng Han, Houfeng Wang, Lizi Liao

Learn the Directions, Normalize the Gains: Post-Training Normalization for LoRA

Read the original on arXiv Machine Learning →

The paper introduces LoRA‑Norm, a post‑training normalization technique for Low‑Rank Adaptation (LoRA) that rebalances the gains of learned singular directions without altering the directions themselves. LoRA‑Norm uses spectral rebalancing and nuclear‑norm restoration to preserve total spectral mass, requiring no calibration data or extra training and adding no inference overhead. Experiments on two backbones and three adaptation tasks show that LoRA‑Norm improves both specialization and capability retention, outperforming other post‑hoc spectral pruning and gradient‑guided editing methods.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 1

Normalized Low-Rank Adaptation

Normalized Low-Rank Adaptation (NoRA) is a lightweight enhancement to the widely used LoRA technique that normalizes the down‑projection matrices during training. By doing so, NoRA stabilizes early optimization dynamics, accelerates convergence, and improves performance across pretraining, supervised fine‑tuning, and reinforcement learning. The method adds no extra trainable parameters or inference‑time cost, making it broadly applicable.

By Jiale Kang, Ziyin Yue, Zheng Zhan, Yangyi Huang, Weiyang Liu
arXiv Machine Learning
Sep 14

Rank-Efficient LoRA via Joint Tangent-Space Optimization under Isotropic Curvature

The paper introduces ISO-LoRA, an optimizer that improves rank utilization in Low‑Rank Adaptation (LoRA) by coupling factor updates through spectral descent on the induced tangent perturbation in weight space. Experiments on GPT‑2 adaptation show that standard optimizers like AdamW concentrate updates in a few singular directions, whereas ISO-LoRA distributes energy more evenly, leading to higher effective rank and better downstream performance across 0.1B‑7B models. The authors provide theoretical guarantees under a stylized spiked‑gradient model and demonstrate that ISO-LoRA consistently outperforms factor‑wise optimizers, especially at moderate‑to‑large LoRA ranks.

By Zihan Zhu, Zhehang Du, Xuyang Chen, Tim Tsz-Kit Lau, Jiayuan Wu, X. Y. Han, Qi Long, Weijie Su