Don't Waste the Noise: Importance-Guided Perturbation Allocation under Joint Global and Local Constraints
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper introduces the Threat Conditional Network (TCN), a model that achieves robust performance across a continuous range of adversarial threat levels. TCN splits representation learning into a threat‑invariant backbone and a lightweight threat‑conditional adaptor, using Fourier‑based embeddings and channel‑wise affine modulation to condition on perturbation budgets. Experiments on CIFAR‑10, CIFAR‑100, and Tiny‑ImageNet demonstrate that TCN matches or exceeds ensembles of budget‑specialized models while adding only 4.6% more parameters, and it generalizes to unseen budgets and mismatched threat conditions.
arXiv:2607. 19855v1 Announce Type: new Abstract: Adversarial robustness is commonly evaluated with predefined attack ensembles, such as AutoAttack, at a single perturbation budget $\varepsilon$ and on a selective choice of perturbation norms.
arXiv:2607. 06109v1 Announce Type: cross Abstract: Multi-perturbation adversarial training (MAT) aims to achieve robustness against multiple $\ell_p$ perturbations but suffers from robustness trade-offs between different threats.
The paper introduces Adversarial Importance Sampling (Advis), a technique that leverages importance sampling over standard training trajectories to estimate and optimize worst‑case returns without extra environment interactions or auxiliary networks, thereby capturing long‑term robustness. It also presents advrl, a modular PyTorch library that consolidates existing robustness methods and adversarial attacks into single‑file implementations for easier prototyping and reproducible evaluation. Finally, the authors highlight that optimal adversarial hyperparameters do not transfer across agents, prompting evaluation against a broader set of attackers (6–14× more configurations) and demonstrate the effectiveness of their approach on continuous control tasks.
arXiv:2606. 11409v1 Announce Type: cross Abstract: Adversarial robustness evaluations of large language models (LLMs) typically report attack success rate (ASR) under fixed query budgets, implicitly treating all attacks as equally costly.
arXiv:2606. 14865v1 Announce Type: cross Abstract: Adversarial Training (AT) improves neural network robustness, but most methods train a fixed parameter space from the start.