OpenAI Blog

Testing robustness against unforeseen adversaries

We’ve developed a method to assess whether a neural network classifier can reliably defend against adversarial attacks not seen during training. Our method yields a new metric, UAR (Unforeseen Attack Robustness), which evaluates the robustness of a single model against an unanticipated attack, and highlights the need to measure performance across a more diverse range of unforeseen attacks.

arXiv Machine Learning
Jun 26

Over-parameterization and Adversarial Robustness in Neural Networks: An Overview and Empirical Analysis

arXiv:2406. 10090v3 Announce Type: replace Abstract: Thanks to their extensive capacity, over-parameterized neural networks exhibit superior predictive capabilities and generalization.

By Srishti Gupta, Zhang Chen, Luca Demetrio, Fabio Brau, Xiaoyi Feng, Zhaoqiang Xia, Antonio Emanuele Cin\`a, Maura Pintor, Luca Oneto, Ambra Demontis, Battista Biggio, Fabio Roli
arXiv Machine Learning
Aug 27

Adversarial Training of Linear Models under Stealthy Attacks

The paper introduces a detector‑based switched model to defend linear predictive models against stealthy false data injection attacks. It derives a convex formulation of the adversarial risk that incorporates protected features and a hyperparameter for attack probability, allowing an explicit trade‑off between clean and attacked data performance. Numerical experiments on real and synthetic datasets demonstrate improved performance on partially attacked data, even when the attack probability is misspecified.

By Lovisa Eriksson, Dave Zachariah, Andr\'e M. H. Teixeira