arXiv AI
Sep 18

Information-Geometric Inverse Distillation for Enhancing Adversarial Transferability

The paper introduces Inverse Knowledge Distillation (IKD), an attack‑agnostic technique that enhances adversarial transferability by maximizing the discrepancy between benign and adversarial prediction distributions on a surrogate model. IKD employs a CE/KL‑equivalent soft‑label objective to push adversarial predictions away from a fixed benign anchor, leveraging Fisher‑sensitive surrogate directions. The authors provide theoretical analysis showing CE and KL induce identical gradients, derive a lower bound on Fisher‑subspace overlap, and demonstrate through extensive ImageNet experiments that IKD consistently improves black‑box attack performance across CNN, ViT, and defended models.

By Wenyuan Wu, Yuan Sun, Yingke Chen, Chao Su, Xi Peng, Dezhong Peng, Xu Wang
arXiv AI
4d ago

Can Pixels Alone Reveal Image Origin? Minimax Limits and Learnable Interfaces for Passive Provenance

The paper investigates whether pixels alone can determine an image’s origin—human, AI class, or specific generator—under adversarial edits. It establishes a minimax limit: the best possible robust acceptance gap equals the minimum total‑variation distance between the target distribution and attacked source distributions, independent of verifier design. The study also shows that practical public verifiers can fail before reaching this theoretical ceiling, highlighting the need to evaluate both statistical limits and deployed verifier behavior separately.

By Kai Yao
arXiv Computer Vision
Sep 18

Fast Preemptive Robustification: High-Frequency Response Anti-Aligns Shared Vulnerability

The paper introduces Fast Preemptive Robustification (FPR), a lightweight defense that enhances the robustness of deep neural networks against transferable adversarial attacks. By sharpening Laplacian responses through a single 3×3 channel‑wise convolution, FPR eliminates the need for surrogate models, iterative optimization, or specialized training. Experiments show that FPR lowers untargeted attack success rates by 12.7% and reduces targeted attack success from 10.7% to 4.1%.

By Jiaming Liang, Chi-Man Pun