arXiv Machine Learning

Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting

arXiv:2508. 04227v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs), spanning predictive architectures to generative Multimodal Large Language Models (MLLMs), have revolutionized artificial intelligence through powerful cross-modal alignment and zero-shot generalization.

arXiv AI
Jul 15

Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning

arXiv:2607. 12112v1 Announce Type: cross Abstract: Federated fine-tuning of Multimodal Large Language Models (MLLMs) across distributed networks enables privacy-sensitive adaptation to evolving data streams, yet a fundamental obstacle prevents robust deployment in dynamic environments: catastrophic forgetting, wherein sequential task updates erase previously acquired knowledge across visual, linguistic, and cross-modal representations.

By Jing Liu, Chenxuanyin Zou, Jiayang Ren, Gaoyun Fang, Chengfang Li, Yan Wang, Zhenchao Ma, Bo Hu
arXiv AI
6d ago

Prompt-Based Continual Compositional Zero-Shot Learning

The paper introduces PromptCCZSL, a framework that enables vision‑language models to continually learn new attributes, objects, and their unique compositions while avoiding forgetting. It uses a frozen VLM backbone with prompt‑based techniques, recency‑weighted multi‑teacher distillation, and several loss functions (CAL, OPL, IDL) to maintain prior knowledge and promote diverse, distinct embeddings. Experiments on UT‑Zappos and C‑GQA show significant performance gains over existing VLM‑based and non‑VLM baselines, establishing a new benchmark for continual compositional zero‑shot learning.

By Sauda Maryam, Sara Nadeem, Faisal Qureshi, Mohsen Ali
arXiv Computer Vision
Aug 28

LLaVAFlow: Preserving Latent Alignment Flow for Parameter-Efficient Multimodal Fine-Tuning

LLaVAFlow is an information‑theoretic distillation framework designed to preserve cross‑modal alignment in Multimodal Large Language Models during visual instruction tuning. It compresses the mutual information between extracted relations and MLLM embeddings to refine alignment flow, and then maximizes mutual information between pretrained and fine‑tuned alignment flows to transfer compact alignment information. Experiments demonstrate that LLaVAFlow effectively maintains alignment flow, improving downstream performance and generalization.

By Muyao Yuan, Muyan Jiao, Jiangyong Ying, Weizhan Zhang, Yuanhong Zhang, Lan Ma, Yuan Gao, Haipeng Du
arXiv Machine Learning
Jul 31

Continual Learning with Vision-Language Models via Semantic-Geometry Preservation

arXiv:2603. 12055v3 Announce Type: replace-cross Abstract: Continual learning of pretrained vision-language models (VLMs) is prone to catastrophic forgetting, yet current approaches adapt to new tasks without explicitly preserving the cross-modal semantic geometry inherited from pretraining and previous stages, allowing new-task supervision to induce geometric distortion.

By Chiyuan He, Zihuan Qiu, Fanman Meng, Runtong Zhang, Linfeng Xu, Qingbo Wu, Hongliang Li
arXiv AI
Jun 19

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models

arXiv:2510. 21978v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has delivered impressive gains in mathematical and multimodal reasoning and has become a standard post-training paradigm for contemporary language and vision-language models.

By Hoang Phan, Xianjun Yang, Yuanshun Yao, Jingyu Zhang, Shengjie Bi, Xiaocheng Tang, Madian Khabsa, Lijuan Liu, Deren Lei
arXiv AI
Sep 7

MePo++: Unifying Representation Refinement and Reconciliation for General Continual Learning

MePo++ is a post‑training framework designed for general continual learning (GCL) that unifies representation refinement and reconciliation. It introduces MetaPrep, which enhances representation plasticity via unsupervised meta‑refinement on pseudo continual sequences, and StreamAlign, which maintains stability by reconciling online features with a stable pretrained geometry. Experiments across various pretrained models, datasets, and continual learning baselines show that MePo++ consistently improves performance in PTM‑based GCL.

By Guanglong Sun, Kanglei Zhou, Liyuan Wang, Qi Cheng, Hongwei Yan, Shuang Cui, Hang Su, Jun Zhu, Yi Zhong
arXiv Computer Vision
Aug 21

OrthoSkillVLA: Continual Skill Learning via Gradient-Informed Skill Subspace Adaptation

arXiv:2608. 19589v1 Announce Type: cross Abstract: Pretrained Vision-Language-Action models provide a strong foundation for robot learning, but sequentially adapting them to diverse skills can perturb the representations and velocity mappings used by previous skills, leading to catastrophic forgetting.

By Jiaqi Wang, Zhou Fang, Qiongfeng Shi, Yi Zhou