arXiv AI

PASs-MoE: Mitigating Misaligned Co-drift among Router and Experts via Pathway Activation Subspaces for Continual Learning

arXiv:2601. 13020v2 Announce Type: replace-cross Abstract: Continual instruction tuning (CIT) requires multimodal large language models (MLLMs) to adapt to a stream of tasks without forgetting prior capabilities.

arXiv Machine Learning
Aug 27

SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning

The paper introduces SAME (Stabilized Mixture-of-Experts) to address challenges in Multimodal Continual Instruction Tuning (MCIT) for large language models. SAME mitigates router drift by decomposing routing dynamics into orthogonal subspaces and updating only task-relevant directions, while preventing expert drift through curvature‑aware scaling that uses historical input covariance without rehearsal. The method also employs adaptive expert activation to freeze selected experts during training, reducing redundant computation and cross‑task interference, and demonstrates state‑of‑the‑art performance on a new long‑task‑sequence benchmark.

By Zhen-Hao Xie, Jun-Tao Tang, Yu-Cheng Shi, Han-Jia Ye, De-Chuan Zhan, Da-Wei Zhou
arXiv Computation and Language
Aug 28

CRAM: Centroid-Routing and Adaptive MoE for Multimodal Continual Instruction Tuning

CRAM (Centroid‑Routing and Adaptive MoE) is a method for Multimodal Continual Instruction Tuning that isolates task‑specific patterns into independent modules to reduce catastrophic forgetting. It uses adaptive‑rank instantiation to allocate only the necessary parameters for new tasks, and centroid‑guided routing with an orthogonality penalty to reuse existing experts while preventing interference. Experiments on diverse benchmarks show CRAM outperforms existing approaches.

By Jun-Tao Tang, Zhen-Hao Xie, Yu-Cheng Shi, Da-Wei Zhou
arXiv Machine Learning
5d ago

New LoRA Skills Should Read but Never Write

The paper introduces READ, a method for composing low‑rank adapters (LoRA) in large language models. By rewriting each adapter into a balanced canonical form and enforcing a one‑directional coupling, READ allows new skills to read but never write into the output subspaces of existing skills, eliminating interference. Experiments on four benchmark suites and two model families show that READ consistently outperforms existing baselines, improving SuperGLUE scores by over twenty points and domain suite scores by more than seven points.

By Zeyan Li, Panqi Yang, Qirong Guo, Shengda Zhuo, SIyuan Qiu, Hu Xu, Chun Li, Jianfeng Xu
arXiv AI
6d ago

Estimating and Orthogonalizing Unknown Pre-training Gradients for Continual Fine-tuning of Large Language Models

The paper introduces EoupCT, a framework that estimates and orthogonalizes unknown pre‑training gradients to mitigate catastrophic forgetting during continual fine‑tuning of large language models. It generates pseudo data most susceptible to forgetting using a learnable soft prompt with Gumbel‑Softmax, then jointly optimizes model parameters and the prompt via a first‑order Pareto optimizer to enforce orthogonality between new task updates and the estimated gradients. Experiments on multiple LLMs show that EoupCT preserves both task‑specific performance and the models’ inherent general‑purpose knowledge.

By Bing Wang, Changchun Li, Xin-Qiang Cai, Lin Yuanbo Wu, Ximing Li, Gang Niu, Masashi Sugiyama