CausAdv: A Causal-based Framework for Detecting Adversarial Examples
arXiv:2411. 00839v4 Announce Type: replace-cross Abstract: Deep learning has led to tremendous success in computer vision, largely due to Convolutional Neural Networks (CNNs).
The paper investigates why robustness training reduces superposition in neural networks. It builds on prior work showing that adversarial examples stem from superposition and that adversarial training diminishes it, but offers no mechanistic explanation. The authors provide an empirical account linking the abandonment of non‑robust features during adversarial training to a reduced number of features overall, thereby lowering superposition.
arXiv:2411. 00839v4 Announce Type: replace-cross Abstract: Deep learning has led to tremendous success in computer vision, largely due to Convolutional Neural Networks (CNNs).
arXiv:2510. 11709v2 Announce Type: replace-cross Abstract: Why do adversarial examples exist, and why do they transfer between models?
arXiv:2607. 12354v1 Announce Type: new Abstract: In this paper, we challenge the prevailing view that information dependency (including rote memorization) drives training data exposure to image reconstruction attacks.
arXiv:2606. 26207v1 Announce Type: cross Abstract: Several theoretical works have tried to explain the adversarial vulnerability of deep neural networks through properties of high-dimensional geometry.
arXiv:2609.06862v1 Announce Type: new Abstract: Superposition refers to neural networks representing more features than they have dimensions. It offers a possible explanation for polysemantic neurons...
arXiv:2410. 07719v4 Announce Type: replace Abstract: Despite being widely adopted as a canonical framework for learning robust models, adversarial training suffers from robust overfitting.
The paper introduces probabilistic adversarial training, a method that enhances robustness by reducing the overlap between a distance-based distribution and a victim-classifier-induced distribution. It derives a KL-based lower bound on probabilistic robustness, which serves as a tractable surrogate objective. Experiments confirm that this approach consistently improves probabilistic robustness and can even boost non-probabilistic adversarial training methods through an induced scaling factor.
arXiv:2406. 10090v3 Announce Type: replace Abstract: Thanks to their extensive capacity, over-parameterized neural networks exhibit superior predictive capabilities and generalization.
Adversarial examples are inputs to machine learning models that an attacker has intentionally designed to cause the model to make a mistake; they’re like optical illusions for machines. In this post we’ll show how adversarial examples work across different mediums, and will discuss why securing systems against them can be difficult.
arXiv:2502. 02260v2 Announce Type: replace Abstract: In the past decade, considerable research effort has been devoted to securing machine learning (ML) models that operate in adversarial settings.
arXiv:2607. 09532v1 Announce Type: new Abstract: We show how an adversarial model trainer can plant backdoors in a large class of deep, feedforward neural networks.