arXiv AI

Why Does Robustness Reduce Superposition?

The paper investigates why robustness training reduces superposition in neural networks. It builds on prior work showing that adversarial examples stem from superposition and that adversarial training diminishes it, but offers no mechanistic explanation. The authors provide an empirical account linking the abandonment of non‑robust features during adversarial training to a reduced number of features overall, thereby lowering superposition.

arXiv AI
3d ago

Probabilistic Adversarial Training

The paper introduces probabilistic adversarial training, a method that enhances robustness by reducing the overlap between a distance-based distribution and a victim-classifier-induced distribution. It derives a KL-based lower bound on probabilistic robustness, which serves as a tractable surrogate objective. Experiments confirm that this approach consistently improves probabilistic robustness and can even boost non-probabilistic adversarial training methods through an induced scaling factor.

By Andi Zhang, Xingyu Zhao, Siddartha Khastgir
arXiv Machine Learning
Jun 26

Over-parameterization and Adversarial Robustness in Neural Networks: An Overview and Empirical Analysis

arXiv:2406. 10090v3 Announce Type: replace Abstract: Thanks to their extensive capacity, over-parameterized neural networks exhibit superior predictive capabilities and generalization.

By Srishti Gupta, Zhang Chen, Luca Demetrio, Fabio Brau, Xiaoyi Feng, Zhaoqiang Xia, Antonio Emanuele Cin\`a, Maura Pintor, Luca Oneto, Ambra Demontis, Battista Biggio, Fabio Roli
OpenAI Blog
Feb 24, 2017

Attacking machine learning with adversarial examples

Adversarial examples are inputs to machine learning models that an attacker has intentionally designed to cause the model to make a mistake; they’re like optical illusions for machines. In this post we’ll show how adversarial examples work across different mediums, and will discuss why securing systems against them can be difficult.