arXiv AI By Sunoh Kim, Daeho Um

SS-TPT: Stability and Suitability-Guided Test-Time Prompt Tuning for Adversarially Robust Vision-Language Models

Read the original on arXiv AI →

arXiv:2606. 06943v1 Announce Type: cross Abstract: Vision-language models (VLMs) such as CLIP achieve strong zero-shot recognition but remain highly fragile under adversarial perturbations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 19

Learning What Not to Learn: Adversarial Disentangled Prompt Tuning for Robust Vision-Language Models

The paper introduces ADAPT, an adversarial disentangled prompt tuning framework designed to improve the robustness of vision‑language models. ADAPT employs a dual‑prompt strategy: a target prompt learns robust features while a set of decoy prompts capture pseudo‑robust, non‑generalizable shortcuts. By enforcing orthogonality between target and decoy prompts, the method mitigates robust generalization overfitting and provides a theoretical error bound for unseen classes, leading to significant empirical robustness gains.

By Yang Chen, Zhan Zhuang, Yanbin Wei, Zebin Chen, Hua Liu, Yu Zhang