Vision Transformers (ViTs) increasingly rely on input-adaptive inference, such as token pruning and early halting, to meet energy and latency budgets. This survey examines a recent class of adversarial efficiency degradation attacks that target these mechanisms to increase computation without necessarily degrading accuracy.
The paper evaluates six preprocessing defenses against adversarial attacks on depthwise‑separable CNNs, the dominant architecture in edge vision systems, and finds that these defenses consistently fail to recover clean predictions for such models, whereas a residual architecture shows partial recovery. The study reveals that the same preprocessing steps that break clean predictions leave adversarial predictions largely intact, creating a measurable asymmetry that can be exploited for detection without retraining or architectural changes. It also demonstrates that common image quality metrics do not reliably indicate defense effectiveness, highlighting a methodological gap in current evaluation practices.
By Jannatul Masruk Mukta, Rifa Sanjida, Adrita Rahman Tory, Md. Saifur Rahman, Khondokar Fida Hasan
The paper introduces TRIM, a black‑box defense for backdoor attacks in computer vision models. TRIM identifies and removes malicious trigger regions at inference time using region‑based segmentation, adaptive trigger discovery via inpainting and diffusion, and selective purification, without needing model internals, training data, or clean samples. Experiments on various datasets and trigger types show TRIM reduces attack success rates to as low as 1.16% while maintaining high clean accuracy.
By Ahmed Abdelnaby, Mohamed Elmahallawy
The paper examines how preprocessing defenses, commonly used to protect edge vision systems, perform on depthwise‑separable CNNs versus residual architectures. Six preprocessing methods were tested against adversarial attacks, revealing that depthwise‑separable models consistently fail to recover from perturbations while residual models show partial recovery. Interestingly, the same preprocessing that hinders clean predictions leaves adversarial predictions largely intact, offering a measurable detection signal, and the study also finds that typical image‑quality metrics do not reliably indicate defense success.
arXiv:2609.39134v1 Announce Type: new
Abstract: Visual-token compression improves the efficiency of large vision-language models, but can expose failures that full-token evaluation misses. We study a...
By Shilinlu Yan, Bowen Chen, Yuechen Zhang, Zhenhong Zhou, Li Sun, Sen Su
arXiv:2607. 02819v1 Announce Type: cross Abstract: Cloud-edge Large Vision-Language Model (LVLM) inference enables efficient deployment by splitting computation between edge devices and cloud servers.
By Zikai Zhang, Rui Hu, Olivera Kotevska, Jiahao Xu
arXiv:2607. 00174v1 Announce Type: cross Abstract: We present a black-box model-stealing attack that recovers private vision-tokenizer configurations of deployed vision-language models (VLMs), including the visual patch size and input preprocessing pipeline.
By Kai Hu, Akash Bharadwaj, Weichen Yu, Matt Fredrikson
arXiv:2607. 25479v1 Announce Type: cross Abstract: Vision--Language Models (VLMs) are increasingly deployed through a model supply chain in which pretrained checkpoints, architecture definitions, text encoders, and exported computation graphs are distributed by third parties and reused across downstream services.
By Maria Rosaria Briglia, Igor Maljkovic, Antonio Emanuele Cin\`a, Luca Oneto, Iacopo Masi, Fabio Roli
arXiv:2606. 04317v1 Announce Type: cross Abstract: Deep neural networks are increasingly deployed across heterogeneous and partially untrusted environments, where models are distributed through cloud storage, CI/CD pipelines, containerized services, and edge execution platforms.
By Bin Duan, Zeyu Bai, Guowei Yang
TEE-X is a TEE‑aware acceleration framework designed to run large vision models, such as Vision Transformers, entirely within Trusted Execution Environments. It introduces a sensitivity‑aware modularization technique and vectorization to overcome memory constraints and latency challenges on edge devices. The framework is validated on OP‑TEE for Arm TrustZone and optimized for the NVIDIA Jetson AGX Xavier, achieving GPU‑level inference latency with minimal accuracy‑latency trade‑offs.
By Kurt M Wilson, Mohaiminul Al Nahian, Abeer Matar A. Almalky, Sadat Shahriyar, Souvik Kundu, Zhishan Guo, Abdullah Al Arafat, Adnan Siraj Rakin
arXiv:2609.05889v1 Announce Type: new
Abstract: Resource-exhaustion attacks against autoregressive vision-language models (VLMs) typically assume unimodal threat models, treating the image branch as...
By Zhaoxiong Ni, Yatie Xiao, Chi-Man Pun, Fei Peng, Qingxiao Guan, Keke Tang
arXiv:2608. 06674v1 Announce Type: cross Abstract: Adversarial vulnerabilities remain a major concern for the safe deployment of neural networks, particularly in object detection, a core task embedded in many safety-critical systems.
By Ridma Jayasundara, Shaheer Mohamed, Tharindu Fernando, Harshala Gammulle, Basura Fernando, Sanka Rasnayake, A V Subramanyam, Sridha Sridharan, Clinton Fookes