arXiv AI

Detect and Suppress: A Mechanistic Defense against Adversarial Patches in VLA Models

The paper presents a mechanistic defense for Vision‑Language‑Action (VLA) models against adversarial patches. By using a sparse autoencoder, the authors identify a feature whose activation correlates strongly with the presence of an adversarial patch and suppress this feature only when a linear probe detects an attack. This conditional intervention improves robustness on the LIBERO‑10 benchmark while avoiding the performance degradation that occurs with continuous suppression.

arXiv Computer Vision
Sep 18

Beyond Patch Removal: Persistent Adversarial Effects in Vision-Language-Action Policies

Adversarial patches applied to Vision‑Language‑Action (VLA) policies not only corrupt actions immediately but also leave lasting state effects that persist after the patch is removed. The study introduces a state‑restoration protocol that evaluates recoverability after patch removal, distinguishing true adversarial impact from occlusion or action‑error magnitude. Experiments on OpenVLA‑OFT with EDPA attacks show that only 36.2% of episodes recover after five chunks, while controls recover at much higher rates; a recovery adapter can improve recovery but its effectiveness drops with delayed intervention.

By Enhao Wu, Fusen Guo, Yuxin Cao, Ziyang Lyu, Lin Li, Wei Song
arXiv AI
Aug 20

Breaking the weakest link to evade vision language models

The paper investigates how Vision Language Models (VLMs) can be fooled by small, human‑imperceptible changes to images. It introduces a gradient‑based attack that targets only the vision encoder, reducing computational cost while still effectively disrupting both untargeted and targeted multimodal alignment. Experiments on open‑source VLMs such as Qwen2.5‑VL, Granite‑Vision, FastVLM, and Phi‑3.5‑Vision demonstrate that these perturbations can significantly alter the models’ textual outputs.

By Ilan Zini, Boussad Addad, Katarzyna Kapusta
Hugging Face Trending Papers
Aug 19

Breaking the weakest link to evade vision language models

The paper investigates how Vision Language Models (VLMs) can be fooled by tiny, human‑imperceptible changes to images. It introduces a gradient‑based attack that targets only the vision encoder, reducing computational cost while still effectively disrupting both untargeted and targeted multimodal interpretations. Experiments on open‑source VLMs such as Qwen2.5‑VL, Granite‑Vision, FastVLM, and Phi‑3.5‑Vision demonstrate that these small perturbations can dramatically alter the models’ textual outputs.

Hugging Face Trending Papers
Sep 17

Beyond Patch Removal: Persistent Adversarial Effects in Vision-Language-Action Policies

The paper investigates how adversarial patches affect Vision‑Language‑Action (VLA) policies, revealing that such patches can cause immediate action corruption and persistent state effects that linger after the patch is removed. A state‑restoration protocol is introduced to isolate these effects by removing the patch at action‑chunk boundaries and measuring recoverability within the remaining step budget. Experiments on OpenVLA-OFT with EDPA attacks show that only 36.2% of episodes recover after five chunks, whereas controls recover at 89.9% and 87.0%. A recovery adapter trained on attack‑induced states improves recovery from 7.7% to 47.4% at one‑chunk latency, but its effectiveness drops sharply with delayed intervention, underscoring the importance of timely recovery.