arXiv AI

Hyperbolic Multimodal Continual Learning: A Closest-Admissible Solution

The paper introduces Hyperbolic Multimodal Continual Learning (HMCL), a method that preserves the Lorentz geometry of hyperbolic multimodal models during sequential updates. By restricting all modalities to a shared hyperbolic isometry, HMCL formulates a joint closest‑admissible (CA) correction—along with a minimal‑rotation (MR) variant—to adjust AdamW updates while maintaining task performance. Experiments on a 16‑task classification‑retrieval stream with three hyperbolic backbones show that HMCL-CA achieves the highest overall score, reduces geometric drift by up to 95.5 %, and improves semantic hierarchy preservation on ImageNet‑WordNet. whyItMatters":"The study demonstrates that explicitly maintaining hyperbolic geometry during continual learning yields superior performance and reduced representation drift compared to existing baselines."

arXiv Machine Learning
Aug 11

Hyperbolic Multimodal Continual Learning

arXiv:2608. 09572v1 Announce Type: new Abstract: Hyperbolic geometry has recently emerged as a powerful representation space for multimodal learning, as it naturally captures hierarchical semantic structure across modalities.

By Jiahong Liu, Ming Shen, Xiaohao Liu, Rex Ying, Menglin Yang, Tat-Seng Chua, Irwin King
arXiv Machine Learning
Jul 31

Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting

arXiv:2508. 04227v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs), spanning predictive architectures to generative Multimodal Large Language Models (MLLMs), have revolutionized artificial intelligence through powerful cross-modal alignment and zero-shot generalization.

By Yuyang Liu, Qiuhe Hong, Linlan Huang, Alexandra Gomez-Villa, Dipam Goswami, Tiantian Peng, Xialei Liu, Joost van de Weijer, Yonghong Tian
arXiv Machine Learning
Jul 31

Continual Learning with Vision-Language Models via Semantic-Geometry Preservation

arXiv:2603. 12055v3 Announce Type: replace-cross Abstract: Continual learning of pretrained vision-language models (VLMs) is prone to catastrophic forgetting, yet current approaches adapt to new tasks without explicitly preserving the cross-modal semantic geometry inherited from pretraining and previous stages, allowing new-task supervision to induce geometric distortion.

By Chiyuan He, Zihuan Qiu, Fanman Meng, Runtong Zhang, Linfeng Xu, Qingbo Wu, Hongliang Li
arXiv Machine Learning
Jun 4

Hyper-ICL: Attention Calibration with Hyperbolic Anchor Distillation for Multimodal In-Context Learning

arXiv:2606. 04434v1 Announce Type: cross Abstract: Multimodal In-Context Learning (ICL) has emerged as a practical inference paradigm for Multimodal Large Language Models, where a small set of interleaved image-text In-Context Demonstrations (ICDs) conditions the model to solve new tasks.

By Niloufar Alipour Talemi, Hossein Kashiani, Fatemeh Afghah
arXiv AI
Jun 8

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models

arXiv:2602. 07026v3 Announce Type: replace-cross Abstract: Despite the success of multimodal contrastive learning in aligning visual and linguistic representations, a persistent geometric anomaly, the Modality Gap, remains: embeddings of distinct modalities expressing identical semantics occupy systematically offset regions.

By Xiaomin Yu, Yi Xin, Yuhui Zhang, Wenjie Zhang, Chonghan Liu, Hanzhen Zhao, Chen Liu, Xiaoxing Hu, Ziyue Qiao, Hao Tang, Xiaobin Hu, Chengwei Qin, Hui Xiong, Yu Qiao, Shuicheng Yan