arXiv AI

SS-TPT: Stability and Suitability-Guided Test-Time Prompt Tuning for Adversarially Robust Vision-Language Models

arXiv:2606. 06943v1 Announce Type: cross Abstract: Vision-language models (VLMs) such as CLIP achieve strong zero-shot recognition but remain highly fragile under adversarial perturbations.

arXiv AI
Aug 19

Learning What Not to Learn: Adversarial Disentangled Prompt Tuning for Robust Vision-Language Models

The paper introduces ADAPT, an adversarial disentangled prompt tuning framework designed to improve the robustness of vision‑language models. ADAPT employs a dual‑prompt strategy: a target prompt learns robust features while a set of decoy prompts capture pseudo‑robust, non‑generalizable shortcuts. By enforcing orthogonality between target and decoy prompts, the method mitigates robust generalization overfitting and provides a theoretical error bound for unseen classes, leading to significant empirical robustness gains.

By Yang Chen, Zhan Zhuang, Yanbin Wei, Zebin Chen, Hua Liu, Yu Zhang
arXiv Computer Vision
Sep 3

Test-Time Logit Prompting for Source-Free Missing Modality Adaptation

The paper introduces Test-Time Logit Prompting (TLP), a lightweight framework that adapts vision-language models to missing-modality inputs without accessing source training data. TLP optimizes logit prompts using uncertainty-aware adjustments and modality-complete consistency regularization, thereby maintaining prediction confidence and semantic consistency. Experiments on various benchmarks show that TLP improves recognition performance by up to 8% while requiring only a few hundred tunable parameters and minimal test-time optimization steps.

By Taixi Chen, Nancy Guo
Hugging Face Trending Papers
Jun 21

Reliability-Guided Adaptive Ensembling for Robust Test-Time Adaptation

Test-time adaptation (TTA) can mitigate domain shift without source data, but it is highly brittle under adversarially contaminated test streams, where corrupted inputs also destabilize online updates. We study robust test-time adaptation (RTTA) in the adversarial-stream setting, which remains comparatively underexplored relative to standard TTA, and propose SAFER (Stochastic Augmentation Framework for Enhanced Robustness), a training-free reliability-guided augmentation wrapper for RTTA.