OpenAI Blog

Transfer of adversarial robustness between perturbation types

arXiv Machine Learning
Jun 26

Over-parameterization and Adversarial Robustness in Neural Networks: An Overview and Empirical Analysis

arXiv:2406. 10090v3 Announce Type: replace Abstract: Thanks to their extensive capacity, over-parameterized neural networks exhibit superior predictive capabilities and generalization.

By Srishti Gupta, Zhang Chen, Luca Demetrio, Fabio Brau, Xiaoyi Feng, Zhaoqiang Xia, Antonio Emanuele Cin\`a, Maura Pintor, Luca Oneto, Ambra Demontis, Battista Biggio, Fabio Roli
arXiv AI
Aug 25

Why Does Robustness Reduce Superposition?

The paper investigates why robustness training reduces superposition in neural networks. It builds on prior work showing that adversarial examples stem from superposition and that adversarial training diminishes it, but offers no mechanistic explanation. The authors provide an empirical account linking the abandonment of non‑robust features during adversarial training to a reduced number of features overall, thereby lowering superposition.

By Adam Elimadi