The Scissors Effect: When Resize-Based Input Diversity Helps or Hurts Transfer Attacks
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2509. 23689v2 Announce Type: replace Abstract: Model Merging (MM) has proven to be an effective alternative to multi-task learning, where several fine-tuned models are merged, without access to the tasks' training data, into one model that retains performance across different tasks.
The paper introduces Inverse Knowledge Distillation (IKD), an attack‑agnostic technique that enhances adversarial transferability by maximizing the discrepancy between benign and adversarial prediction distributions on a surrogate model. IKD employs a CE/KL‑equivalent soft‑label objective to push adversarial predictions away from a fixed benign anchor, leveraging Fisher‑sensitive surrogate directions. The authors provide theoretical analysis showing CE and KL induce identical gradients, derive a lower bound on Fisher‑subspace overlap, and demonstrate through extensive ImageNet experiments that IKD consistently improves black‑box attack performance across CNN, ViT, and defended models.
arXiv:2609.10002v1 Announce Type: new Abstract: Deepfake detectors remain vulnerable to transfer-based black-box attacks, in which adversarial examples are generated on a source surrogate model and t...
The paper investigates whether pixels alone can determine an image’s origin—human, AI class, or specific generator—under adversarial edits. It establishes a minimax limit: the best possible robust acceptance gap equals the minimum total‑variation distance between the target distribution and attacked source distributions, independent of verifier design. The study also shows that practical public verifiers can fail before reaching this theoretical ceiling, highlighting the need to evaluate both statistical limits and deployed verifier behavior separately.
arXiv:2605.25663v2 Announce Type: replace-cross Abstract: Black-box adversarial attacks that minimize only the ground-truth confidence suffer from class drift: perturbations wander through the featur...
The paper introduces Fast Preemptive Robustification (FPR), a lightweight defense that enhances the robustness of deep neural networks against transferable adversarial attacks. By sharpening Laplacian responses through a single 3×3 channel‑wise convolution, FPR eliminates the need for surrogate models, iterative optimization, or specialized training. Experiments show that FPR lowers untargeted attack success rates by 12.7% and reduces targeted attack success from 10.7% to 4.1%.