arXiv Computer Vision By Enhao Wu, Fusen Guo, Yuxin Cao, Ziyang Lyu, Lin Li, Wei Song

Beyond Patch Removal: Persistent Adversarial Effects in Vision-Language-Action Policies

Read the original on arXiv Computer Vision →

Adversarial patches applied to Vision‑Language‑Action (VLA) policies not only corrupt actions immediately but also leave lasting state effects that persist after the patch is removed. The study introduces a state‑restoration protocol that evaluates recoverability after patch removal, distinguishing true adversarial impact from occlusion or action‑error magnitude. Experiments on OpenVLA‑OFT with EDPA attacks show that only 36.2% of episodes recover after five chunks, while controls recover at much higher rates; a recovery adapter can improve recovery but its effectiveness drops with delayed intervention.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

Hugging Face Trending Papers
Sep 17

Beyond Patch Removal: Persistent Adversarial Effects in Vision-Language-Action Policies

The paper investigates how adversarial patches affect Vision‑Language‑Action (VLA) policies, revealing that such patches can cause immediate action corruption and persistent state effects that linger after the patch is removed. A state‑restoration protocol is introduced to isolate these effects by removing the patch at action‑chunk boundaries and measuring recoverability within the remaining step budget. Experiments on OpenVLA-OFT with EDPA attacks show that only 36.2% of episodes recover after five chunks, whereas controls recover at 89.9% and 87.0%. A recovery adapter trained on attack‑induced states improves recovery from 7.7% to 47.4% at one‑chunk latency, but its effectiveness drops sharply with delayed intervention, underscoring the importance of timely recovery.

arXiv AI
Jun 11

Diffusion-based Cumulative Adversarial Purification for Vision Language Models

arXiv:2506. 03933v2 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) have shown remarkable capabilities in multimodal understanding, yet their susceptibility to adversarial perturbations poses a significant threat to their reliability in real-world applications.

By Jia Fu, Yongtao Wu, Yihang Chen, Kunyu Peng, Xiao Zhang, Volkan Cevher, Sepideh Pashami, Anders Holst