arXiv Computer Vision

Capacity Overflow: A Blind Spot for Backdoor Attacks in Vision MoE

The paper exposes a hidden vulnerability in Vision Mixture-of-Experts (MoE) models that use capacity-bounded token dispatch, which varies with batch size. It presents a three-phase backdoor attack: injecting a backdoor into an early MoE layer, training a neutralizer in a deeper layer to suppress it under normal capacity, and then adjusting the batch-adaptive capacity factor so that the neutralizer is disabled when large batches are used at deployment. Experiments on V-MoE and Swin-MoE show high attack success rates (76‑87%) for large batches while keeping the attack dormant and undetected during small-batch audits, evading several state‑of‑the‑art defenses.

Hugging Face Trending Papers
Aug 5

A Survey of Adversarial Efficiency Degradation for Vision Transformer by Exploiting Input-adaptive Optimization

Vision Transformers (ViTs) increasingly rely on input-adaptive inference, such as token pruning and early halting, to meet energy and latency budgets. This survey examines a recent class of adversarial efficiency degradation attacks that target these mechanisms to increase computation without necessarily degrading accuracy.

arXiv Computer Vision
Sep 4

Preprocessing Failure and Adversarial Detection in Depthwise-Separable Edge Vision Systems

The paper evaluates six preprocessing defenses against adversarial attacks on depthwise‑separable CNNs, the dominant architecture in edge vision systems, and finds that these defenses consistently fail to recover clean predictions for such models, whereas a residual architecture shows partial recovery. The study reveals that the same preprocessing steps that break clean predictions leave adversarial predictions largely intact, creating a measurable asymmetry that can be exploited for detection without retraining or architectural changes. It also demonstrates that common image quality metrics do not reliably indicate defense effectiveness, highlighting a methodological gap in current evaluation practices.

By Jannatul Masruk Mukta, Rifa Sanjida, Adrita Rahman Tory, Md. Saifur Rahman, Khondokar Fida Hasan
arXiv Computer Vision
Sep 4

Beyond Small Patches: Black-Box Detection and Purification of Diverse Backdoor Triggers

The paper introduces TRIM, a black‑box defense for backdoor attacks in computer vision models. TRIM identifies and removes malicious trigger regions at inference time using region‑based segmentation, adaptive trigger discovery via inpainting and diffusion, and selective purification, without needing model internals, training data, or clean samples. Experiments on various datasets and trigger types show TRIM reduces attack success rates to as low as 1.16% while maintaining high clean accuracy.

By Ahmed Abdelnaby, Mohamed Elmahallawy
Hugging Face Trending Papers
Sep 3

Preprocessing Failure and Adversarial Detection in Depthwise-Separable Edge Vision Systems

The paper examines how preprocessing defenses, commonly used to protect edge vision systems, perform on depthwise‑separable CNNs versus residual architectures. Six preprocessing methods were tested against adversarial attacks, revealing that depthwise‑separable models consistently fail to recover from perturbations while residual models show partial recovery. Interestingly, the same preprocessing that hinders clean predictions leaves adversarial predictions largely intact, offering a measurable detection signal, and the study also finds that typical image‑quality metrics do not reliably indicate defense success.

arXiv AI
Jul 29

Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering

arXiv:2607. 25479v1 Announce Type: cross Abstract: Vision--Language Models (VLMs) are increasingly deployed through a model supply chain in which pretrained checkpoints, architecture definitions, text encoders, and exported computation graphs are distributed by third parties and reused across downstream services.

By Maria Rosaria Briglia, Igor Maljkovic, Antonio Emanuele Cin\`a, Luca Oneto, Iacopo Masi, Fabio Roli
arXiv AI
Aug 25

TEE-X: TEE-aware Acceleration Framework for Large Vision Models at the Edge

TEE-X is a TEE‑aware acceleration framework designed to run large vision models, such as Vision Transformers, entirely within Trusted Execution Environments. It introduces a sensitivity‑aware modularization technique and vectorization to overcome memory constraints and latency challenges on edge devices. The framework is validated on OP‑TEE for Arm TrustZone and optimized for the NVIDIA Jetson AGX Xavier, achieving GPU‑level inference latency with minimal accuracy‑latency trade‑offs.

By Kurt M Wilson, Mohaiminul Al Nahian, Abeer Matar A. Almalky, Sadat Shahriyar, Souvik Kundu, Zhishan Guo, Abdullah Al Arafat, Adnan Siraj Rakin
arXiv Machine Learning
Aug 10

Corrupting Attention: Evasion-Based Adversarial Attacks on Encoder Attention in Detection Transformers

arXiv:2608. 06674v1 Announce Type: cross Abstract: Adversarial vulnerabilities remain a major concern for the safe deployment of neural networks, particularly in object detection, a core task embedded in many safety-critical systems.

By Ridma Jayasundara, Shaheer Mohamed, Tharindu Fernando, Harshala Gammulle, Basura Fernando, Sanka Rasnayake, A V Subramanyam, Sridha Sridharan, Clinton Fookes