arXiv AI

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models

arXiv:2607. 02819v1 Announce Type: cross Abstract: Cloud-edge Large Vision-Language Model (LVLM) inference enables efficient deployment by splitting computation between edge devices and cloud servers.

arXiv AI
Sep 2

Position Matters: Feature Inversion Attacks in ViT Split Inference with Token Reduction and Shuffling

The paper investigates privacy risks in Vision Transformer (ViT) split‑inference systems that use token reduction and token shuffling to lower computation and communication costs. It shows that even after token shuffling, transmitted token embeddings still contain enough positional information for a new attack, the Spatially Aligned Reconstruction Attack (SARA), which predicts token positions, restores spatial layout, fills missing embeddings with a masked autoencoder, and reconstructs the input image. While token reduction offers stronger protection, significant leakage remains when retained tokens preserve semantic and positional cues, and the authors propose a lightweight edge‑side defense that removes positional embeddings and adapts transformer blocks via knowledge distillation to reduce SARA’s effectiveness without harming downstream accuracy.

By Stefano Leggio, Giulio Rossolini, Alessandro Biondi
arXiv AI
Jul 29

Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering

arXiv:2607. 25479v1 Announce Type: cross Abstract: Vision--Language Models (VLMs) are increasingly deployed through a model supply chain in which pretrained checkpoints, architecture definitions, text encoders, and exported computation graphs are distributed by third parties and reused across downstream services.

By Maria Rosaria Briglia, Igor Maljkovic, Antonio Emanuele Cin\`a, Luca Oneto, Iacopo Masi, Fabio Roli
arXiv Computer Vision
Aug 27

Capacity Overflow: A Blind Spot for Backdoor Attacks in Vision MoE

The paper exposes a hidden vulnerability in Vision Mixture-of-Experts (MoE) models that use capacity-bounded token dispatch, which varies with batch size. It presents a three-phase backdoor attack: injecting a backdoor into an early MoE layer, training a neutralizer in a deeper layer to suppress it under normal capacity, and then adjusting the batch-adaptive capacity factor so that the neutralizer is disabled when large batches are used at deployment. Experiments on V-MoE and Swin-MoE show high attack success rates (76‑87%) for large batches while keeping the attack dormant and undetected during small-batch audits, evading several state‑of‑the‑art defenses.

By Xiaocheng Zou, Tiancheng Zheng, Xiaolin Xu, Ruyi Ding
Hugging Face Trending Papers
Aug 5

A Survey of Adversarial Efficiency Degradation for Vision Transformer by Exploiting Input-adaptive Optimization

Vision Transformers (ViTs) increasingly rely on input-adaptive inference, such as token pruning and early halting, to meet energy and latency budgets. This survey examines a recent class of adversarial efficiency degradation attacks that target these mechanisms to increase computation without necessarily degrading accuracy.

Hugging Face Trending Papers
Aug 19

Breaking the weakest link to evade vision language models

The paper investigates how Vision Language Models (VLMs) can be fooled by tiny, human‑imperceptible changes to images. It introduces a gradient‑based attack that targets only the vision encoder, reducing computational cost while still effectively disrupting both untargeted and targeted multimodal interpretations. Experiments on open‑source VLMs such as Qwen2.5‑VL, Granite‑Vision, FastVLM, and Phi‑3.5‑Vision demonstrate that these small perturbations can dramatically alter the models’ textual outputs.

arXiv AI
Aug 20

Breaking the weakest link to evade vision language models

The paper investigates how Vision Language Models (VLMs) can be fooled by small, human‑imperceptible changes to images. It introduces a gradient‑based attack that targets only the vision encoder, reducing computational cost while still effectively disrupting both untargeted and targeted multimodal alignment. Experiments on open‑source VLMs such as Qwen2.5‑VL, Granite‑Vision, FastVLM, and Phi‑3.5‑Vision demonstrate that these perturbations can significantly alter the models’ textual outputs.

By Ilan Zini, Boussad Addad, Katarzyna Kapusta