arXiv AI

Probabilistic Adversarial Training

The paper introduces probabilistic adversarial training, a method that enhances robustness by reducing the overlap between a distance-based distribution and a victim-classifier-induced distribution. It derives a KL-based lower bound on probabilistic robustness, which serves as a tractable surrogate objective. Experiments confirm that this approach consistently improves probabilistic robustness and can even boost non-probabilistic adversarial training methods through an induced scaling factor.

arXiv AI
Sep 18

Information-Geometric Inverse Distillation for Enhancing Adversarial Transferability

The paper introduces Inverse Knowledge Distillation (IKD), an attack‑agnostic technique that enhances adversarial transferability by maximizing the discrepancy between benign and adversarial prediction distributions on a surrogate model. IKD employs a CE/KL‑equivalent soft‑label objective to push adversarial predictions away from a fixed benign anchor, leveraging Fisher‑sensitive surrogate directions. The authors provide theoretical analysis showing CE and KL induce identical gradients, derive a lower bound on Fisher‑subspace overlap, and demonstrate through extensive ImageNet experiments that IKD consistently improves black‑box attack performance across CNN, ViT, and defended models.

By Wenyuan Wu, Yuan Sun, Yingke Chen, Chao Su, Xi Peng, Dezhong Peng, Xu Wang