arXiv AI
Jul 15

Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning

arXiv:2607. 12112v1 Announce Type: cross Abstract: Federated fine-tuning of Multimodal Large Language Models (MLLMs) across distributed networks enables privacy-sensitive adaptation to evolving data streams, yet a fundamental obstacle prevents robust deployment in dynamic environments: catastrophic forgetting, wherein sequential task updates erase previously acquired knowledge across visual, linguistic, and cross-modal representations.

By Jing Liu, Chenxuanyin Zou, Jiayang Ren, Gaoyun Fang, Chengfang Li, Yan Wang, Zhenchao Ma, Bo Hu
arXiv Computer Vision
Aug 27

A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training

The paper introduces a new task called Multimodal Unsupervised Continual Post-Training (MU‑CPT), which allows multimodal large language models (MLLMs) to continuously learn from streaming unlabeled data. It identifies token‑level visual dependence (VD) as essential for MU‑CPT, using its structural distortion to detect cross‑modal forgetting and its heterogeneity to guide new‑task learning. The proposed Visual Dependence‑Aware (VDA) framework includes Visually Constrained Optimal Transport (VC‑OT) to mitigate forgetting and Visually Modulated Adaptation (VMA) to enhance new‑task plasticity, achieving a balance between stability and adaptability in MU‑CPT.

By Kaichen Li, Zhilin Zhu, Jianhao Huang, Zhengqin Lai, Baochen Xiong, Zibo Shao, Yaguang Song, Linhui Xiao, Xiaoshan Yang, Changsheng Xu