arXiv:2604.15678v2 Announce Type: replace
Abstract: Pretrained Vision-Language Models (VLMs) like CLIP show promise in continual learning, but existing Few-Shot Class-Incremental Learning (FSCIL) met...
By Eunju Lee, MiHyeon Kim, JuneHyoung Kwon, Yoonji Lee, JiHyun Kim, Soojin Jang, YoungBin Kim
The paper investigates why reducing the modality gap between image and text representations in CLIP does not always improve zero‑shot classification accuracy. It shows that while average alignment improves, the relative decision margins among classes can shift, leading to a prediction‑level hubness where predictions concentrate on a few classes. Experiments across datasets confirm that accuracy drops correlate with increased prediction concentration for both linear and learning‑based gap corrections.
By Shota Sato, Hajime Kiyama, Tosho Hirasawa, Mamoru Komachi
arXiv:2609.06967v1 Announce Type: cross
Abstract: Ensuring effective transfer learning for vision-language models without compromising their generalization performance is crucial. However, many exist...
By Seungmin Oh, Seunghun Kang, Jongbin Ryu
arXiv:2602. 21397v2 Announce Type: replace-cross Abstract: Prompt learning has become a dominant paradigm for adapting vision-language models (VLMs) such as CLIP to downstream tasks without modifying pretrained weights.
By Sajjad Ghiasvand, Haniyeh Ehsani Oskouie, Mahnoosh Alizadeh, Ramtin Pedarsani
arXiv:2608.29395v1 Announce Type: new
Abstract: Vision-language models such as CLIP and SigLIP provide strong zero-shot recognition, but their predictions can degrade when deployed on target data tha...
By Pedram MohajerAnsari, Amir Salarpour, Run Wang, Mert D. Pes\'e
The paper introduces UnInfo, a test‑time adaptation method for vision‑language models like CLIP that addresses image corruption—a realistic distribution shift caused by sensor conditions. UnInfo leverages uniformity‑aware confidence maximization, information‑aware loss balancing, and knowledge distillation from an EMA teacher to preserve embedding uniformity and improve zero‑shot classification accuracy. Experiments show that UnInfo outperforms existing TTA methods on corrupted image datasets.
By Kazuki Adachi, Shin'ya Yamaguchi, Tomoki Hamagami
The paper introduces Domain Recentering with Confidence Calibration (DRC), a training‑free technique that adapts CLIP to unlabeled target images by fitting a Gaussian mixture and subtracting a posterior‑weighted average of component means from each embedding. It further corrects residual class bias using a log‑prior adjustment based on confidence‑weighted predictions. DRC outperforms other methods, raising average accuracy on cross‑domain datasets by 4.13 and 5.07 points over zero‑shot CLIP for ViT‑B/16 and ResNet‑50, and maintains gains under ImageNet distribution shifts.
By Youngeun Seol, Jimin Shin, Heeseo Yoon, Uiwon Hwang
arXiv:2608. 18339v1 Announce Type: cross Abstract: Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts during inference.
By Qi Yu, Zhichen Zeng, Katherine Tieu, Xiyuan Yang, Ruizhong Qiu, Yuchen Yan, Lihui Liu, Yanjun Zhao, Lingjie Chen, Jingrui He, Hanghang Tong
The paper introduces Domain Recentering with Confidence Calibration (DRC), a training‑free technique that adapts CLIP to unlabeled target images by fitting a Gaussian mixture and applying posterior‑weighted mean subtraction, followed by a log‑prior correction based on confidence‑weighted predictions. DRC improves cross‑domain accuracy, surpassing zero‑shot CLIP by 4.13 and 5.07 points on ViT‑B/16 and ResNet‑50, and maintains gains under ImageNet distribution shifts.
arXiv:2609.24564v1 Announce Type: new
Abstract: CLIP, a foundational vision-language model, has emerged as a powerful tool for open-vocabulary semantic segmentation. While freezing CLIP's text encode...
By Zelin Peng, Zhengqin Xu, Changsong Wen, Yu Huang, Yaoming Wang, Xiaokang Yang, Wei Shen
arXiv:2512.04305v3 Announce Type: replace
Abstract: Vision-language models (VLMs) such as CLIP are increasingly adapted across decentralized data silos, yet the reliability of their predictions under...
By Mainak Singha, Masih Aminbeidokhti, Paolo Casari, Gianni Franchi, Elisa Ricci, Subhankar Roy
arXiv:2609.15640v1 Announce Type: cross
Abstract: Long-text image--text congruence scoring is increasingly important for vision-language systems that must evaluate whether detailed textual descriptio...
By Alessandro Gambetti, Qiwei Han