arXiv Computer Vision

Continual Visual Learning under Evolving Semantic Concept Shift

arXiv Machine Learning
Jul 31

Continual Learning with Vision-Language Models via Semantic-Geometry Preservation

arXiv:2603. 12055v3 Announce Type: replace-cross Abstract: Continual learning of pretrained vision-language models (VLMs) is prone to catastrophic forgetting, yet current approaches adapt to new tasks without explicitly preserving the cross-modal semantic geometry inherited from pretraining and previous stages, allowing new-task supervision to induce geometric distortion.

By Chiyuan He, Zihuan Qiu, Fanman Meng, Runtong Zhang, Linfeng Xu, Qingbo Wu, Hongliang Li
arXiv Computer Vision
Aug 31

Dual-Stream Semantic Guidance with Prototype Anchor Calibration for Source-Fully-Free Adaptation of Vision-Language Models

The paper introduces Dual-Stream Semantic Guidance (DSSG), a framework for Source‑Fully‑Free Domain Adaptation of Vision‑Language Models that mitigates dual semantic drift through a caption stream and a class‑anchor stream. It adds a Dynamic Cross‑Modal Knowledge Distillation module and a Prototype Anchor Calibration extension (DSSG‑PAC) to reduce computation while maintaining performance. Experiments show DSSG outperforms state‑of‑the‑art methods and DSSG‑PAC cuts adaptation time by 18.9% with minimal loss in accuracy.

By Weiwei Xiang, Shun Peng, Guangyi Xiao, Hao Chen, Lei Yang
arXiv Computer Vision
Aug 28

When Semantics Saturate or Emerge: Adaptation-Conditional Semantic Utility in Source-Free Cross-Domain Few-Shot Learning

The paper investigates whether language prompts selected by zero‑shot accuracy remain effective after visual adaptation in source‑free cross‑domain few‑shot learning. Using a paired protocol, the authors compare generic class‑name templates with detailed class descriptions before and after Low‑Rank Adaptation (LoRA) on datasets such as EuroSAT, CropDisease, ISIC, and ChestX. They identify two regimes: semantic saturation, where detailed prompts are already useful before adaptation, and semantic emergence, where detailed prompts become more useful only after visual representation updates, driven by changes in prediction patterns.

By Wei Liu, Xing Deng, Haijian Shao
arXiv Computation and Language
Aug 31

The Telephone Game: Evaluating Semantic Drift in Unified Models

The paper introduces the Semantic Drift Protocol (SDP), a Telephone Game-inspired method to evaluate how well unified models preserve meaning when alternating between image-to-text (I2T) and text-to-image (T2I) generation over multiple generations. It defines Mean Cumulative Drift (MCD) and Multi-Generation GenEval (MGG) as metrics for semantic retention, and presents a new benchmark of 400 image‑text pairs from NoCaps and DOCCI to stress-test models beyond COCO. Applying SDP to seven models shows that models with strong single‑pass scores can still suffer severe semantic drift, revealing catastrophic failure modes that isolated benchmarks miss.

By Sabbir Mollah, Rohit Gupta, Sirnam Swetha, Qingyang Liu, Ahnaf Munir, Mubarak Shah
arXiv AI
Aug 5

Does Forgetting Transfer Across Modalities? A Real-World Benchmark for Cross-Modal Knowledge Unlearning Evaluation

arXiv:2608. 03791v1 Announce Type: new Abstract: Vision-Language Models (VLMs), like Large Language Models (LLMs), may memorize sensitive, copyrighted, or harmful knowledge from their pretraining corpora.

By Chunlin Liu, Junnian Chen, Haitong Jiang, Jianyu Zhao, Yingsen Pang, Jingchen Li, Jiabiao He, Youming Lu, Jinhe Bi, Yuntao Du
arXiv Computer Vision
Aug 26

AffineTok: Semantic Affine Consistency for Diffusion-Friendly Visual Tokenizer

arXiv:2608.23864v1 Announce Type: new Abstract: Visual tokenizers increasingly inject semantic supervision into latent spaces to make downstream diffusion easier. Yet how these semantics should be or...

By Junqiu Yu, Pandeng Li, Yikai Wang, Jiaxing Zhao, Yujie Wei, Kaixun Jiang, Quanhao Li, Hongtao Yu, Zhihang Liu, Zhaohe Liao, Junjie Zhou, Yun Zheng, Yu Liu, Yanwei Fu
arXiv Computer Vision
Aug 27

A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training

The paper introduces a new task called Multimodal Unsupervised Continual Post-Training (MU‑CPT), which allows multimodal large language models (MLLMs) to continuously learn from streaming unlabeled data. It identifies token‑level visual dependence (VD) as essential for MU‑CPT, using its structural distortion to detect cross‑modal forgetting and its heterogeneity to guide new‑task learning. The proposed Visual Dependence‑Aware (VDA) framework includes Visually Constrained Optimal Transport (VC‑OT) to mitigate forgetting and Visually Modulated Adaptation (VMA) to enhance new‑task plasticity, achieving a balance between stability and adaptability in MU‑CPT.

By Kaichen Li, Zhilin Zhu, Jianhao Huang, Zhengqin Lai, Baochen Xiong, Zibo Shao, Yaguang Song, Linhui Xiao, Xiaoshan Yang, Changsheng Xu