Statistical Adversaries: Natural Backdoor-like Adversarial Features in Clean Vision Datasets
arXiv:2607. 05516v2 Announce Type: replace-cross Abstract: Model-specific adversarial attacks have been extensively studied.
arXiv:2411. 00839v4 Announce Type: replace-cross Abstract: Deep learning has led to tremendous success in computer vision, largely due to Convolutional Neural Networks (CNNs).
arXiv:2607. 05516v2 Announce Type: replace-cross Abstract: Model-specific adversarial attacks have been extensively studied.
arXiv:2607. 05516v1 Announce Type: cross Abstract: Model-specific adversarial attacks have been extensively studied.
arXiv:2606. 02267v1 Announce Type: new Abstract: The vulnerability of deep neural networks to adversarial examples poses a significant challenge for real-world deployment.
Model-specific adversarial attacks have been extensively studied. We study a different failure mode: naturally occurring statistical signals in vision data that can behave like backdoor-like triggers without being maliciously inserted.
The paper introduces a new framework for generating medical image counterfactuals that does not rely on auxiliary generative models. By extracting causal evidence directly from a classifier, the method deterministically produces edits within user-specified regions, requiring no additional training. Experiments on real-world medical imaging datasets show that these counterfactuals alter classifier predictions while staying closer to the original image than generative baselines, offering a clearer view of the model’s decision boundary.
arXiv:2412. 08394v2 Announce Type: replace Abstract: Deep neural networks (DNNs) are vulnerable to adversarial samples crafted by adding imperceptible perturbations to clean data, potentially leading to incorrect and dangerous predictions.
arXiv:2506. 03933v2 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) have shown remarkable capabilities in multimodal understanding, yet their susceptibility to adversarial perturbations poses a significant threat to their reliability in real-world applications.
arXiv:2504. 14798v2 Announce Type: replace Abstract: Machine Unlearning (MUL) has emerged as a key mechanism for privacy protection and content regulation, yet current techniques often fail to guarantee the complete removal of sensitive information.
The paper introduces a new method for generating medical image counterfactuals that does not rely on auxiliary generative models. By extracting causal evidence directly from the classifier, the approach deterministically edits user-specified regions to alter predictions while staying closer to the original image than generative baselines. Experiments on real-world medical imaging datasets show that this technique provides a more direct and transparent view of the classifier’s decision boundary.
arXiv:2510. 11709v2 Announce Type: replace-cross Abstract: Why do adversarial examples exist, and why do they transfer between models?
arXiv:2511. 13749v2 Announce Type: replace Abstract: Deep neural networks are known to be vulnerable to adversarial perturbations, which are small, carefully crafted inputs that lead to incorrect predictions.
The paper investigates why robustness training reduces superposition in neural networks. It builds on prior work showing that adversarial examples stem from superposition and that adversarial training diminishes it, but offers no mechanistic explanation. The authors provide an empirical account linking the abandonment of non‑robust features during adversarial training to a reduced number of features overall, thereby lowering superposition.