The paper introduces Foundation Preserving LoRA (FoLoRA), a forgetting‑aware optimization framework that balances adaptation to downstream tasks with preservation of pretraining capabilities. FoLoRA uses a first‑order preservation condition to define a forgetting penalty based on pretraining‑proxy activations and a task utility from downstream activations, scoring update directions via a generalized Rayleigh quotient. This spectral coordinate system enables gated Adam updates that reduce low‑utility, high‑penalty directions, and the method constructs pretraining proxy calibration data by sampling from the pretrained model. Experiments on math, code, and instruction‑following tasks demonstrate that FoLoRA achieves a stronger balance between target task performance and aggregate preservation of non‑target capabilities compared to baselines.
By Dongjun Kim, Adrian de Wynter, Huancheng Chen, Heasung Kim, Haris Vikalo
arXiv:2602.03493v2 Announce Type: replace
Abstract: Low-Rank Adaptation (LoRA) methods have emerged as crucial techniques for adapting large pre-trained models to downstream tasks under computational...
By Alessio Quercia, Arya Bangun, Ira Assent, Hanno Scharr
arXiv:2602. 00722v2 Announce Type: replace Abstract: Parameter-efficient continual learning aims to adapt pre-trained models to sequential tasks without forgetting previously acquired knowledge.
By Hao Gu, Mao-Lin Luo, Zi-Hao Zhou, Han-Chen Zhang, Min-Ling Zhang, Tong Wei
The paper introduces EoupCT, a framework that estimates and orthogonalizes unknown pre‑training gradients to mitigate catastrophic forgetting during continual fine‑tuning of large language models. It generates pseudo data most susceptible to forgetting using a learnable soft prompt with Gumbel‑Softmax, then jointly optimizes model parameters and the prompt via a first‑order Pareto optimizer to enforce orthogonality between new task updates and the estimated gradients. Experiments on multiple LLMs show that EoupCT preserves both task‑specific performance and the models’ inherent general‑purpose knowledge.
By Bing Wang, Changchun Li, Xin-Qiang Cai, Lin Yuanbo Wu, Ximing Li, Gang Niu, Masashi Sugiyama
arXiv:2608. 12332v1 Announce Type: cross Abstract: In recent years, low-rank adaptation (LoRA) has emerged as a significant paradigm that freezes pre-trained weights and introduces small, learnable adapters instead of fine-tuning the full set of parameters.
By Hyowon Wi, Noseong Park
arXiv:2610.02126v1 Announce Type: cross
Abstract: We explore catastrophic forgetting in the context of large pre-trained models. By considering forgetting as a geometric problem in the input space of...
By Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes