arXiv AI

Can Pixels Alone Reveal Image Origin? Minimax Limits and Learnable Interfaces for Passive Provenance

The paper investigates whether pixels alone can determine an image’s origin—human, AI class, or specific generator—under adversarial edits. It establishes a minimax limit: the best possible robust acceptance gap equals the minimum total‑variation distance between the target distribution and attacked source distributions, independent of verifier design. The study also shows that practical public verifiers can fail before reaching this theoretical ceiling, highlighting the need to evaluate both statistical limits and deployed verifier behavior separately.

arXiv Machine Learning
Sep 22

Reinforcement Learning Inspired Black-box Adversarial Attacks for Computer Vision

The paper introduces RIBA, a reinforcement‑learning inspired black‑box adversarial attack that generates perturbations for neural networks with fewer queries than existing methods. RIBA achieves a 25.4% reduction in median queries on ResNet‑18/Cifar10 and a 22.5% reduction on Vit‑B/16/ImageNet, while matching white‑box attack performance on an adversarially trained model.

By Florian Krone, Elena Hoemann, Sven Hallerbach
arXiv AI
Aug 14

SPARED: Reasoning-Based AI-Generated Image Detection via Adversarially Edited Data

arXiv:2608. 12876v1 Announce Type: cross Abstract: Detecting AI-generated images is only half the task: a deployed detector must also justify its verdict, yet existing detectors inherit three failure modes from their training data: real and fake images collected from different sources invite provenance shortcuts, supervised explanation corpora teach templated rationales, and a static forgery corpus leaves the decision boundary standing still while generators keep moving.

By Yicheng Bao, Xiahui Guo, Xuhong Wang, Xin Tan
arXiv AI
Aug 11

Attack Ensembles Expose a Safety-Utility Trade-off in Black-Box Guard Defenses Against Encoded VLM Jailbreaks

arXiv:2607. 26574v2 Announce Type: replace-cross Abstract: Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or text rendered inside an image slips past a guard that would block it in plain language - the decode gap.

By Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Yi Feng, Xiao Luo, Zijian Xiao, Haowen Xu, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita
arXiv AI
4d ago

Blind, Not Weak: A Best-of-Suite Safety-Utility Frontier for Recover-and-Reguard Defenses Against Encoded VLM Jailbreaks

The paper presents a preprocessor that recovers and decodes encoded content in vision‑language models to close the decode gap that allows harmful requests to bypass safety classifiers. Evaluated against eleven encoding attacks, the preprocessor raises block rates from 0 % to 67‑90 % but also increases benign over‑refusal, and no configuration achieves an ensemble attack‑success rate below 40 % while keeping benign over‑refusal under 70 %. The study shows that closing one encoding channel merely relocates success rather than eliminating it, highlighting the limits of recovery‑based defenses.

By Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Hanwen Liu, Yi Feng, Haowen Xu, Xiangchen Guan, Yang Chen, Zijian Xiao, Xiao Luo, Mohammad Zandsalimy, Shanu Sushmita