arXiv Machine Learning

Shift-Aware Calibration for Fine-Tuned CLIP: Leveraging Image-Text Alignment

arXiv:2501. 19060v4 Announce Type: replace-cross Abstract: Vision-language models (VLMs), such as CLIP, adapt effectively to downstream tasks through prompt tuning, but fine-tuning can misalign predictive confidence and accuracy, particularly on unseen classes.

arXiv Computation and Language
Sep 2

When Modality Gap Reduction Fails: Prediction-Level Hubness in CLIP

The paper investigates why reducing the modality gap between image and text representations in CLIP does not always improve zero‑shot classification accuracy. It shows that while average alignment improves, the relative decision margins among classes can shift, leading to a prediction‑level hubness where predictions concentrate on a few classes. Experiments across datasets confirm that accuracy drops correlate with increased prediction concentration for both linear and learning‑based gap corrections.

By Shota Sato, Hajime Kiyama, Tosho Hirasawa, Mamoru Komachi
arXiv Computer Vision
Sep 3

Uniformity First: Uniformity-aware Test-time Adaptation of Vision-language Models against Image Corruption

The paper introduces UnInfo, a test‑time adaptation method for vision‑language models like CLIP that addresses image corruption—a realistic distribution shift caused by sensor conditions. UnInfo leverages uniformity‑aware confidence maximization, information‑aware loss balancing, and knowledge distillation from an EMA teacher to preserve embedding uniformity and improve zero‑shot classification accuracy. Experiments show that UnInfo outperforms existing TTA methods on corrupted image datasets.

By Kazuki Adachi, Shin'ya Yamaguchi, Tomoki Hamagami
arXiv AI
Sep 25

Domain Recentering and Confidence-Weighted Prior Calibration for Vision-Language Models

The paper introduces Domain Recentering with Confidence Calibration (DRC), a training‑free technique that adapts CLIP to unlabeled target images by fitting a Gaussian mixture and subtracting a posterior‑weighted average of component means from each embedding. It further corrects residual class bias using a log‑prior adjustment based on confidence‑weighted predictions. DRC outperforms other methods, raising average accuracy on cross‑domain datasets by 4.13 and 5.07 points over zero‑shot CLIP for ViT‑B/16 and ResNet‑50, and maintains gains under ImageNet distribution shifts.

By Youngeun Seol, Jimin Shin, Heeseo Yoon, Uiwon Hwang
arXiv AI
Aug 20

From Inference to Adaptation: A Unified Optimal Transport View of Vision Language Model

arXiv:2608. 18339v1 Announce Type: cross Abstract: Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts during inference.

By Qi Yu, Zhichen Zeng, Katherine Tieu, Xiyuan Yang, Ruizhong Qiu, Yuchen Yan, Lihui Liu, Yanjun Zhao, Lingjie Chen, Jingrui He, Hanghang Tong
Hugging Face Trending Papers
Sep 24

Domain Recentering and Confidence-Weighted Prior Calibration for Vision-Language Models

The paper introduces Domain Recentering with Confidence Calibration (DRC), a training‑free technique that adapts CLIP to unlabeled target images by fitting a Gaussian mixture and applying posterior‑weighted mean subtraction, followed by a log‑prior correction based on confidence‑weighted predictions. DRC improves cross‑domain accuracy, surpassing zero‑shot CLIP by 4.13 and 5.07 points on ViT‑B/16 and ResNet‑50, and maintains gains under ImageNet distribution shifts.