USE: A Unified Self-Ensembling Framework for Test-Time Prompt Tuning
arXiv:2607. 03900v1 Announce Type: cross Abstract: Test-time adaptation (TTA) has emerged as a popular paradigm for improving the performance of vision-language models (e.
arXiv:2606. 14299v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) such as CLIP have become a standard backbone for open-vocabulary recognition, yet their zero-shot predictions remain vulnerable to distribution shifts encountered at deployment.
arXiv:2607. 03900v1 Announce Type: cross Abstract: Test-time adaptation (TTA) has emerged as a popular paradigm for improving the performance of vision-language models (e.
The paper introduces selective adaptation for vision‑language models, questioning whether test‑time adaptation (TTA) should always be applied. By analyzing per‑sample predictions before and after adaptation, the authors find that many adaptations are negligible or even harmful, flipping correct predictions. They propose Cross‑Augmentation Similarity (CAS), which skips adaptation when predictions across augmented views are highly similar, achieving comparable or better accuracy while reducing adaptation by up to 85%.
arXiv:2608.29395v1 Announce Type: new Abstract: Vision-language models such as CLIP and SigLIP provide strong zero-shot recognition, but their predictions can degrade when deployed on target data tha...
The paper introduces SubTTA, a test-time adaptation method for vision‑language models that aligns the semantic subspaces of visual and textual modalities to improve zero‑shot predictions. It addresses two issues: the modality gap caused by distribution shifts and visual nuisance that masks task‑specific semantics. By minimizing chordal distance between principal subspaces and projecting visual features onto a task‑specific textual subspace, SubTTA refines decision boundaries and achieves an average 2.24% improvement over existing TTA methods.
arXiv:2606. 06943v1 Announce Type: cross Abstract: Vision-language models (VLMs) such as CLIP achieve strong zero-shot recognition but remain highly fragile under adversarial perturbations.
arXiv:2605. 03403v2 Announce Type: replace-cross Abstract: Group Relative Policy Optimization (GRPO) has recently shown strong performance in post-training large language models and vision-language models.
arXiv:2606. 28551v1 Announce Type: cross Abstract: Building performant Vision-Language Models (VLMs) requires carefully curating large-scale training datasets, yet the community lacks systematic benchmarks for evaluating such curation strategies.
The paper surveys Continual Test-Time Adaptation (CTTA), a framework that adapts pretrained computer‑vision models to non‑stationary target distributions without source data or labeled targets, while avoiding catastrophic forgetting and error accumulation. It formally defines the CTTA problem, categorizes existing methods into optimization‑based, parameter‑efficient, and architecture‑based families, and reviews representative techniques and benchmarks across standard evaluation settings. The survey also outlines current limitations and proposes future research directions, such as adapting foundation models and black‑box systems.
The paper introduces FairTPT, a fairness-aware test‑time prompt tuning method for vision‑language models like CLIP. It jointly minimizes target marginal entropy while maximizing spurious marginal entropy to reduce bias under subpopulation shifts. Experiments show that standard episodic test‑time adaptation can worsen disparities, but FairTPT outperforms existing debiasing methods while preserving overall performance.
The paper presents Test-Time Adaptation via Cache Personalization (TTA‑CaP), a gradient‑free, cache‑based method that personalizes vision‑language models for facial expression recognition in videos. TTA‑CaP uses three complementary caches—a personalized static cache, a positive target cache, and a negative target cache—controlled by a tri‑gate mechanism to prevent corruption and provide robust subject‑matched evidence. Experiments on BioVid, StressID, and BAH datasets show that TTA‑CaP outperforms state‑of‑the‑art test‑time adaptation methods while keeping computational and memory overhead low.
Deep neural nets achieve remarkable performance when training and test data share the same distribution, but this assumption frequently breaks in real-world deployment, where data undergoes continual distributional shifts. Continual Test-Time Adaptation (CTTA) addresses this challenge by adapting pretrained models to non-stationary target distributions on-the-fly, without access to source data or labeled targets, while mitigating two critical failure modes: catastrophic forgetting of source knowledge and error accumulation from noisy pseudo-labels over extended time horizons.
The paper introduces Test-Time Logit Prompting (TLP), a lightweight framework that adapts vision-language models to missing-modality inputs without accessing source training data. TLP optimizes logit prompts using uncertainty-aware adjustments and modality-complete consistency regularization, thereby maintaining prediction confidence and semantic consistency. Experiments on various benchmarks show that TLP improves recognition performance by up to 8% while requiring only a few hundred tunable parameters and minimal test-time optimization steps.