arXiv Computer Vision

Anchoring Adversarial Trajectories to Data Manifolds: A Bilevel Transfer Optimization Framework

The paper introduces Manifold Anchored Bilevel Transfer (MABT), a framework that aligns adversarial attack trajectories with the intrinsic data manifold to reduce surrogate-specific overfitting. MABT employs a relaxed manifold-anchoring operator as a semantic rectifier and formulates transfer attack generation as a distributional bilevel optimization problem, learning geometry-aligned initializations that minimize expected transfer risk. A Hessian-free solver with linear-time complexity is developed to efficiently solve the resulting hierarchy, and experiments show improved transferability across 10 attackers, 28 configurations, various victim models, and defense mechanisms.

arXiv AI
Sep 18

Information-Geometric Inverse Distillation for Enhancing Adversarial Transferability

The paper introduces Inverse Knowledge Distillation (IKD), an attack‑agnostic technique that enhances adversarial transferability by maximizing the discrepancy between benign and adversarial prediction distributions on a surrogate model. IKD employs a CE/KL‑equivalent soft‑label objective to push adversarial predictions away from a fixed benign anchor, leveraging Fisher‑sensitive surrogate directions. The authors provide theoretical analysis showing CE and KL induce identical gradients, derive a lower bound on Fisher‑subspace overlap, and demonstrate through extensive ImageNet experiments that IKD consistently improves black‑box attack performance across CNN, ViT, and defended models.

By Wenyuan Wu, Yuan Sun, Yingke Chen, Chao Su, Xi Peng, Dezhong Peng, Xu Wang
arXiv Statistics ML
6d ago

Brenier Meets Adversarial Training: Optimal Transport Geometry for Robust Learning

The paper introduces a penalized distributionally robust optimization framework that allows an adversary to choose any distribution while incurring a Wasserstein penalty for deviating from the empirical distribution. It shows that the adversary’s problem can be reformulated as optimizing transport maps that push empirical samples to adversarial ones, proving that optimal maps are cyclically monotone. The authors argue that standard per-sample adversarial training violates this property and propose two remedies—multi-start particle ascent and input-convex neural network parameterization—to enforce cyclical monotonicity, demonstrating improved robustness and generalization in experiments on regression, image classification, and control tasks.

By Alireza Abdollahpoorrostam, Ehsan Sharifian, Buse \c{S}en, Marco Cuturi, Daniel Kuhn
arXiv Machine Learning
Sep 1

ARMOR: Manifold-Oriented Training for Adversarially Robust Aerial Object Detection under Data Scarcity

arXiv:2608.29510v1 Announce Type: cross Abstract: Aerial object detection is increasingly deployed in real-world applications, but models remain vulnerable to physical, universal adversarial patches...

By Haoran Wang, Matthew Lau, Alec Helbling, Matthew Hull, ShengYun Peng, Mansi Phute, Martin Andreoni, Willian T. Lunardi, Duen Horng Chau, Wenke Lee
arXiv Machine Learning
Aug 27

Adversarial Training of Linear Models under Stealthy Attacks

The paper introduces a detector‑based switched model to defend linear predictive models against stealthy false data injection attacks. It derives a convex formulation of the adversarial risk that incorporates protected features and a hyperparameter for attack probability, allowing an explicit trade‑off between clean and attacked data performance. Numerical experiments on real and synthetic datasets demonstrate improved performance on partially attacked data, even when the attack probability is misspecified.

By Lovisa Eriksson, Dave Zachariah, Andr\'e M. H. Teixeira
arXiv Machine Learning
5d ago

Frame the adversary: a structure-aware attack methodology

The paper introduces a new framework for creating frequency‑based adversarial attacks that are grounded in an explicit optimization problem. By defining a perturbation constraint set linked to structured, non‑orthogonal transforms, the authors show that attacks can be generated as weighted σ‒projections onto this set, providing a clear geometric characterization. Experiments on standard datasets demonstrate that these attacks are highly effective across both pretrained and robust models, even on unseen architectures.

By Vicky Kouni, Stelios Perrakis, Francis Bach, Pascal Frossard, Yann Chevaleyre