The paper introduces Foundation Preserving LoRA (FoLoRA), a forgetting‑aware optimization framework that balances adaptation to downstream tasks with preservation of pretraining capabilities. FoLoRA uses a first‑order preservation condition to define a forgetting penalty based on pretraining‑proxy activations and a task utility from downstream activations, scoring update directions via a generalized Rayleigh quotient. This spectral coordinate system enables gated Adam updates that reduce low‑utility, high‑penalty directions, and the method constructs pretraining proxy calibration data by sampling from the pretrained model. Experiments on math, code, and instruction‑following tasks demonstrate that FoLoRA achieves a stronger balance between target task performance and aggregate preservation of non‑target capabilities compared to baselines.
By Dongjun Kim, Adrian de Wynter, Huancheng Chen, Heasung Kim, Haris Vikalo
arXiv:2602.03493v2 Announce Type: replace
Abstract: Low-Rank Adaptation (LoRA) methods have emerged as crucial techniques for adapting large pre-trained models to downstream tasks under computational...
By Alessio Quercia, Arya Bangun, Ira Assent, Hanno Scharr
arXiv:2602. 00722v2 Announce Type: replace Abstract: Parameter-efficient continual learning aims to adapt pre-trained models to sequential tasks without forgetting previously acquired knowledge.
By Hao Gu, Mao-Lin Luo, Zi-Hao Zhou, Han-Chen Zhang, Min-Ling Zhang, Tong Wei
The paper introduces EoupCT, a framework that estimates and orthogonalizes unknown pre‑training gradients to mitigate catastrophic forgetting during continual fine‑tuning of large language models. It generates pseudo data most susceptible to forgetting using a learnable soft prompt with Gumbel‑Softmax, then jointly optimizes model parameters and the prompt via a first‑order Pareto optimizer to enforce orthogonality between new task updates and the estimated gradients. Experiments on multiple LLMs show that EoupCT preserves both task‑specific performance and the models’ inherent general‑purpose knowledge.
By Bing Wang, Changchun Li, Xin-Qiang Cai, Lin Yuanbo Wu, Ximing Li, Gang Niu, Masashi Sugiyama
arXiv:2608. 12332v1 Announce Type: cross Abstract: In recent years, low-rank adaptation (LoRA) has emerged as a significant paradigm that freezes pre-trained weights and introduces small, learnable adapters instead of fine-tuning the full set of parameters.
By Hyowon Wi, Noseong Park
arXiv:2610.02126v1 Announce Type: cross
Abstract: We explore catastrophic forgetting in the context of large pre-trained models. By considering forgetting as a geometric problem in the input space of...
By Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes
arXiv:2601. 13020v2 Announce Type: replace-cross Abstract: Continual instruction tuning (CIT) requires multimodal large language models (MLLMs) to adapt to a stream of tasks without forgetting prior capabilities.
By Zhiyan Hou, Haiyun Guo, Haokai Ma, Yandu Sun, Yonghui Yang, Jinqiao Wang
Normalized Low-Rank Adaptation (NoRA) is a lightweight enhancement to the widely used LoRA technique that normalizes the down‑projection matrices during training. By doing so, NoRA stabilizes early optimization dynamics, accelerates convergence, and improves performance across pretraining, supervised fine‑tuning, and reinforcement learning. The method adds no extra trainable parameters or inference‑time cost, making it broadly applicable.
By Jiale Kang, Ziyin Yue, Zheng Zhan, Yangyi Huang, Weiyang Liu
arXiv:2510. 18874v3 Announce Type: replace Abstract: Adapting language models (LMs) to new tasks via post-training carries the risk of degrading existing capabilities -- a phenomenon classically known as catastrophic forgetting.
By Howard Chen, Noam Razin, Karthik Narasimhan, Danqi Chen
arXiv:2607. 09202v1 Announce Type: cross Abstract: Continual learning commonly relies on post-hoc mechanisms such as replay, elastic regularization, or distillation.
By Julius St\"ork
arXiv:2606. 23487v2 Announce Type: replace Abstract: Medical vision-language models (VLMs) such as BiomedCLIP generalize broadly, but adapting them to a clinical service is as much a safety problem as an accuracy one.
By Rishabh Jha, Amrita Singh, Prashanna Chudal
arXiv:2608. 16249v1 Announce Type: new Abstract: Machine unlearning in Large Language Models (LLMs) faces a critical trade-off between erasing target knowledge and preserving general utility.
By Jaewan Choi, Junyoung Yang, Sangdon Park