arXiv:2605. 20247v2 Announce Type: replace-cross Abstract: Catastrophic forgetting remains a major obstacle to continual learning in large language models (LLMs) and vision--language models (VLMs).
By Yang Liu, Toan Nguyen, Flora D. Salim
The paper introduces SAME (Stabilized Mixture-of-Experts) to address challenges in Multimodal Continual Instruction Tuning (MCIT) for large language models. SAME mitigates router drift by decomposing routing dynamics into orthogonal subspaces and updating only task-relevant directions, while preventing expert drift through curvature‑aware scaling that uses historical input covariance without rehearsal. The method also employs adaptive expert activation to freeze selected experts during training, reducing redundant computation and cross‑task interference, and demonstrates state‑of‑the‑art performance on a new long‑task‑sequence benchmark.
By Zhen-Hao Xie, Jun-Tao Tang, Yu-Cheng Shi, Han-Jia Ye, De-Chuan Zhan, Da-Wei Zhou
CRAM (Centroid‑Routing and Adaptive MoE) is a method for Multimodal Continual Instruction Tuning that isolates task‑specific patterns into independent modules to reduce catastrophic forgetting. It uses adaptive‑rank instantiation to allocate only the necessary parameters for new tasks, and centroid‑guided routing with an orthogonality penalty to reuse existing experts while preventing interference. Experiments on diverse benchmarks show CRAM outperforms existing approaches.
By Jun-Tao Tang, Zhen-Hao Xie, Yu-Cheng Shi, Da-Wei Zhou
arXiv:2607. 23837v1 Announce Type: new Abstract: Large language models generalize well to individual tasks but lack an inherent mechanism for learning them sequentially, leading to catastrophic forgetting.
By Reza Rahimi Azghan, Gautham Krishna Gudur, Giulia Pedrielli, Pavan Turaga, Hassan Ghasemzadeh
arXiv:2606. 10338v1 Announce Type: cross Abstract: Machine unlearning is increasingly important for large language models, yet unlearning in Mixture-of-Experts (MoE) architectures remains underexplored.
By Jingyi Xie, Yijun Lin, Yinjiang Xiong, Zhikun Zhang, Sai Li
arXiv:2601. 18699v2 Announce Type: replace Abstract: Sequential fine-tuning of Large Language Models (LLMs) adaptation to target tasks often triggers catastrophic forgetting, where the acquisition of novel target skills degrades ancestral capabilities.
By Gustav Olaf Yunus Laitinen-Fredriksson Lundstrom-Imanov
arXiv:2606. 07500v1 Announce Type: cross Abstract: Continual learning in Large Language Models (LLMs) is hindered by the plasticity-stability dilemma, where acquiring new capabilities often leads to catastrophic forgetting of previous knowledge.
By Fatema Siddika, Md Anwar Hossen, Tanwi Mallick, Ali Jannesari
arXiv:2605.07111v3 Announce Type: replace-cross
Abstract: Recent literature on fine-tuning Large Language Models highlights a fundamental debate. While Full Fine-Tuning (FFT) provides greater represe...
By Haozhan Tang, Xiuqi Zhu, Xinyin Zhang, Boxun Li, Virginia Smith, Kevin Kuo
The paper introduces READ, a method for composing low‑rank adapters (LoRA) in large language models. By rewriting each adapter into a balanced canonical form and enforcing a one‑directional coupling, READ allows new skills to read but never write into the output subspaces of existing skills, eliminating interference. Experiments on four benchmark suites and two model families show that READ consistently outperforms existing baselines, improving SuperGLUE scores by over twenty points and domain suite scores by more than seven points.
By Zeyan Li, Panqi Yang, Qirong Guo, Shengda Zhuo, SIyuan Qiu, Hu Xu, Chun Li, Jianfeng Xu
arXiv:2606. 28117v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) has become the standard tool for parameter-efficient fine-tuning of large pretrained models.
By Tanguy Dieudonn\'e, Giulia Lanzillotta, Enis Simsar, Louis Barinka, Thomas Hofmann
arXiv:2604.23036v2 Announce Type: replace-cross
Abstract: Despite MoE models leading many benchmarks, supervised fine-tuning (SFT) for the MoE architectures remains difficult because its router layer...
By Haoze He, Xingyuan Ding, Xuan Jiang, Xinkai Zou, Alex Cheng, Yibo Zhao, Juncheng Billy Li, Heather Miller
The paper introduces EoupCT, a framework that estimates and orthogonalizes unknown pre‑training gradients to mitigate catastrophic forgetting during continual fine‑tuning of large language models. It generates pseudo data most susceptible to forgetting using a learnable soft prompt with Gumbel‑Softmax, then jointly optimizes model parameters and the prompt via a first‑order Pareto optimizer to enforce orthogonality between new task updates and the estimated gradients. Experiments on multiple LLMs show that EoupCT preserves both task‑specific performance and the models’ inherent general‑purpose knowledge.
By Bing Wang, Changchun Li, Xin-Qiang Cai, Lin Yuanbo Wu, Ximing Li, Gang Niu, Masashi Sugiyama