arXiv Machine Learning

Learning with Bilevel-Minimax Optimization for Efficient and Reliable Transfer Attacks

arXiv:2608. 11815v1 Announce Type: new Abstract: Transfer-based adversarial attacks craft adversarial examples using surrogate models to mislead black-box victim models.

arXiv Computer Vision
3d ago

Anchoring Adversarial Trajectories to Data Manifolds: A Bilevel Transfer Optimization Framework

The paper introduces Manifold Anchored Bilevel Transfer (MABT), a framework that aligns adversarial attack trajectories with the intrinsic data manifold to reduce surrogate-specific overfitting. MABT employs a relaxed manifold-anchoring operator as a semantic rectifier and formulates transfer attack generation as a distributional bilevel optimization problem, learning geometry-aligned initializations that minimize expected transfer risk. A Hessian-free solver with linear-time complexity is developed to efficiently solve the resulting hierarchy, and experiments show improved transferability across 10 attackers, 28 configurations, various victim models, and defense mechanisms.

By Yaohua Liu, Yifan Guo, Jiaxin Gao
arXiv AI
Sep 18

Information-Geometric Inverse Distillation for Enhancing Adversarial Transferability

The paper introduces Inverse Knowledge Distillation (IKD), an attack‑agnostic technique that enhances adversarial transferability by maximizing the discrepancy between benign and adversarial prediction distributions on a surrogate model. IKD employs a CE/KL‑equivalent soft‑label objective to push adversarial predictions away from a fixed benign anchor, leveraging Fisher‑sensitive surrogate directions. The authors provide theoretical analysis showing CE and KL induce identical gradients, derive a lower bound on Fisher‑subspace overlap, and demonstrate through extensive ImageNet experiments that IKD consistently improves black‑box attack performance across CNN, ViT, and defended models.

By Wenyuan Wu, Yuan Sun, Yingke Chen, Chao Su, Xi Peng, Dezhong Peng, Xu Wang
arXiv AI
Jun 18

Revealing Hidden Vulnerabilities in Autoencoders through Gradient Signal Restoration

arXiv:2505. 03646v5 Announce Type: replace-cross Abstract: Adversarial robustness of deep autoencoders (AEs) has received less attention than that of discriminative models, although their compressed latent representations induce ill-conditioned mappings that can amplify small input perturbations and destabilize reconstructions.

By Chethan Krishnamurthy Ramanaik, Arjun Roy, Tobias Callies, Eirini Ntoutsi