arXiv:2607. 02020v1 Announce Type: new Abstract: Multimodal large language models must continually adapt to evolving tasks and domains, yet standard continual learning metrics mainly measure whether old answers remain correct, leaving the stability of multimodal grounding largely unexamined.
By Qianyu Chen, Canran Xiao, Runxuan Tang
arXiv:2609.07009v1 Announce Type: new
Abstract: Multimodal continual learning has recently shown great potential for developing agents with human-like intelligence by continuously learning new tasks...
By Kai Guo, Chuanbin Liu, Peng Hu, Hao Wang, Xi Peng
The paper introduces a new task called Multimodal Unsupervised Continual Post-Training (MU‑CPT), which allows multimodal large language models (MLLMs) to continuously learn from streaming unlabeled data. It identifies token‑level visual dependence (VD) as essential for MU‑CPT, using its structural distortion to detect cross‑modal forgetting and its heterogeneity to guide new‑task learning. The proposed Visual Dependence‑Aware (VDA) framework includes Visually Constrained Optimal Transport (VC‑OT) to mitigate forgetting and Visually Modulated Adaptation (VMA) to enhance new‑task plasticity, achieving a balance between stability and adaptability in MU‑CPT.
By Kaichen Li, Zhilin Zhu, Jianhao Huang, Zhengqin Lai, Baochen Xiong, Zibo Shao, Yaguang Song, Linhui Xiao, Xiaoshan Yang, Changsheng Xu
In this paper, we explore a novel task of Multimodal Unsupervised Continual Post-Training (MU-CPT), enabling deployed MLLMs to continually evolve from streaming unlabeled data. Existing unsupervised p...
The paper introduces Hyperbolic Multimodal Continual Learning (HMCL), a method that preserves the Lorentz geometry of hyperbolic multimodal models during sequential updates. By restricting all modalities to a shared hyperbolic isometry, HMCL formulates a joint closest‑admissible (CA) correction—along with a minimal‑rotation (MR) variant—to adjust AdamW updates while maintaining task performance. Experiments on a 16‑task classification‑retrieval stream with three hyperbolic backbones show that HMCL-CA achieves the highest overall score, reduces geometric drift by up to 95.5 %, and improves semantic hierarchy preservation on ImageNet‑WordNet.
whyItMatters":"The study demonstrates that explicitly maintaining hyperbolic geometry during continual learning yields superior performance and reduced representation drift compared to existing baselines."
By Jiahong Liu, Ming Shen, Xiaohao Liu, Rex Ying, Menglin Yang, Tat-Seng Chua, Irwin King
arXiv:2605. 20247v2 Announce Type: replace-cross Abstract: Catastrophic forgetting remains a major obstacle to continual learning in large language models (LLMs) and vision--language models (VLMs).
By Yang Liu, Toan Nguyen, Flora D. Salim