Vision Transformers (ViTs) increasingly rely on input-adaptive inference, such as token pruning and early halting, to meet energy and latency budgets. This survey examines a recent class of adversarial efficiency degradation attacks that target these mechanisms to increase computation without necessarily degrading accuracy.
The paper evaluates six preprocessing defenses against adversarial attacks on depthwise‑separable CNNs, the dominant architecture in edge vision systems, and finds that these defenses consistently fail to recover clean predictions for such models, whereas a residual architecture shows partial recovery. The study reveals that the same preprocessing steps that break clean predictions leave adversarial predictions largely intact, creating a measurable asymmetry that can be exploited for detection without retraining or architectural changes. It also demonstrates that common image quality metrics do not reliably indicate defense effectiveness, highlighting a methodological gap in current evaluation practices.
By Jannatul Masruk Mukta, Rifa Sanjida, Adrita Rahman Tory, Md. Saifur Rahman, Khondokar Fida Hasan
The paper introduces TRIM, a black‑box defense for backdoor attacks in computer vision models. TRIM identifies and removes malicious trigger regions at inference time using region‑based segmentation, adaptive trigger discovery via inpainting and diffusion, and selective purification, without needing model internals, training data, or clean samples. Experiments on various datasets and trigger types show TRIM reduces attack success rates to as low as 1.16% while maintaining high clean accuracy.
By Ahmed Abdelnaby, Mohamed Elmahallawy
The paper examines how preprocessing defenses, commonly used to protect edge vision systems, perform on depthwise‑separable CNNs versus residual architectures. Six preprocessing methods were tested against adversarial attacks, revealing that depthwise‑separable models consistently fail to recover from perturbations while residual models show partial recovery. Interestingly, the same preprocessing that hinders clean predictions leaves adversarial predictions largely intact, offering a measurable detection signal, and the study also finds that typical image‑quality metrics do not reliably indicate defense success.
arXiv:2609.39134v1 Announce Type: new
Abstract: Visual-token compression improves the efficiency of large vision-language models, but can expose failures that full-token evaluation misses. We study a...
By Shilinlu Yan, Bowen Chen, Yuechen Zhang, Zhenhong Zhou, Li Sun, Sen Su
arXiv:2607. 02819v1 Announce Type: cross Abstract: Cloud-edge Large Vision-Language Model (LVLM) inference enables efficient deployment by splitting computation between edge devices and cloud servers.
By Zikai Zhang, Rui Hu, Olivera Kotevska, Jiahao Xu