arXiv Machine Learning By Song-Lin Lv, Yu-Yang Chen, Zhi Zhou, Lan-Zhe Guo

Shift-Aware Calibration for Fine-Tuned CLIP: Leveraging Image-Text Alignment

Read the original on arXiv Machine Learning →

arXiv:2501. 19060v4 Announce Type: replace-cross Abstract: Vision-language models (VLMs), such as CLIP, adapt effectively to downstream tasks through prompt tuning, but fine-tuning can misalign predictive confidence and accuracy, particularly on unseen classes.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 2

When Modality Gap Reduction Fails: Prediction-Level Hubness in CLIP

The paper investigates why reducing the modality gap between image and text representations in CLIP does not always improve zero‑shot classification accuracy. It shows that while average alignment improves, the relative decision margins among classes can shift, leading to a prediction‑level hubness where predictions concentrate on a few classes. Experiments across datasets confirm that accuracy drops correlate with increased prediction concentration for both linear and learning‑based gap corrections.

By Shota Sato, Hajime Kiyama, Tosho Hirasawa, Mamoru Komachi
arXiv Computer Vision
Sep 3

Uniformity First: Uniformity-aware Test-time Adaptation of Vision-language Models against Image Corruption

The paper introduces UnInfo, a test‑time adaptation method for vision‑language models like CLIP that addresses image corruption—a realistic distribution shift caused by sensor conditions. UnInfo leverages uniformity‑aware confidence maximization, information‑aware loss balancing, and knowledge distillation from an EMA teacher to preserve embedding uniformity and improve zero‑shot classification accuracy. Experiments show that UnInfo outperforms existing TTA methods on corrupted image datasets.

By Kazuki Adachi, Shin'ya Yamaguchi, Tomoki Hamagami