A New Kind of Adversarial Example: Measuring the Human-Model Gap, and Its Relationship to OOD Detection
arXiv:2607. 22722v1 Announce Type: cross Abstract: Almost all adversarial attacks add an imperceptible perturbation to fool a model.
The paper investigates whether pixels alone can determine an image’s origin—human, AI class, or specific generator—under adversarial edits. It establishes a minimax limit: the best possible robust acceptance gap equals the minimum total‑variation distance between the target distribution and attacked source distributions, independent of verifier design. The study also shows that practical public verifiers can fail before reaching this theoretical ceiling, highlighting the need to evaluate both statistical limits and deployed verifier behavior separately.
arXiv:2607. 22722v1 Announce Type: cross Abstract: Almost all adversarial attacks add an imperceptible perturbation to fool a model.
arXiv:2609.10002v1 Announce Type: new Abstract: Deepfake detectors remain vulnerable to transfer-based black-box attacks, in which adversarial examples are generated on a source surrogate model and t...
arXiv:2606.22516v2 Announce Type: replace Abstract: Input Diversity (DI), a random resize and pad applied at each attack iteration, is a near-default ingredient of transfer-based attacks, widely assu...
arXiv:2605.25663v2 Announce Type: replace-cross Abstract: Black-box adversarial attacks that minimize only the ground-truth confidence suffer from class drift: perturbations wander through the featur...
The paper introduces RIBA, a reinforcement‑learning inspired black‑box adversarial attack that generates perturbations for neural networks with fewer queries than existing methods. RIBA achieves a 25.4% reduction in median queries on ResNet‑18/Cifar10 and a 22.5% reduction on Vit‑B/16/ImageNet, while matching white‑box attack performance on an adversarially trained model.
arXiv:2608. 15113v1 Announce Type: cross Abstract: Learned image compression (LIC) has demonstrated remarkable rate-distortion (RD) performance in benign settings.
arXiv:2608. 12876v1 Announce Type: cross Abstract: Detecting AI-generated images is only half the task: a deployed detector must also justify its verdict, yet existing detectors inherit three failure modes from their training data: real and fake images collected from different sources invite provenance shortcuts, supervised explanation corpora teach templated rationales, and a static forgery corpus leaves the decision boundary standing still while generators keep moving.
arXiv:2607. 26574v2 Announce Type: replace-cross Abstract: Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or text rendered inside an image slips past a guard that would block it in plain language - the decode gap.
arXiv:2606. 03647v1 Announce Type: cross Abstract: Accurately evaluating adversarial robustness is a longstanding challenge.
arXiv:2607. 19855v1 Announce Type: new Abstract: Adversarial robustness is commonly evaluated with predefined attack ensembles, such as AutoAttack, at a single perturbation budget $\varepsilon$ and on a selective choice of perturbation norms.
arXiv:2609.07147v1 Announce Type: new Abstract: Federated learning, as a privacy-preserving distributed machine learning paradigm, faces significant threats from backdoor attacks. Compared to central...
The paper presents a preprocessor that recovers and decodes encoded content in vision‑language models to close the decode gap that allows harmful requests to bypass safety classifiers. Evaluated against eleven encoding attacks, the preprocessor raises block rates from 0 % to 67‑90 % but also increases benign over‑refusal, and no configuration achieves an ensemble attack‑success rate below 40 % while keeping benign over‑refusal under 70 %. The study shows that closing one encoding channel merely relocates success rather than eliminating it, highlighting the limits of recovery‑based defenses.