arXiv AI By Akchunya Chanchal, David A. Kelly, Hana Chockler

Activation-Deactivation: A General Framework for Robust Post-hoc Explainable AI

Read the original on arXiv AI →

arXiv:2510. 01038v2 Announce Type: replace Abstract: Perturbation-based explainability methods face criticism due to their reliance on out-of-distribution mutants.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
22h ago

Learning What Not to Learn: Adversarial Disentangled Prompt Tuning for Robust Vision-Language Models

arXiv:2608. 17306v1 Announce Type: cross Abstract: While adversarial prompt tuning can enhance robustness of vision-language models efficiently, we find that existing methods aggravate robust generalization overfitting on seen classes, leading to a rapid degradation in performance against adversarial examples of unseen classes as training progresses.

By Yang Chen, Zhan Zhuang, Yanbin Wei, Zebin Chen, Hua Liu, Yu Zhang