Hugging Face Trending Papers

AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models

Read the original on Hugging Face Trending Papers →

Vision-language models (VLMs) are increasingly deployed on infrared (IR) remote sensing imagery in security-critical settings, yet their adversarial robustness remains unexamined. We present AirflowAttack, to our knowledge the first adversarial attack for IR remote-sensing VLMs and the first to weaponize thermal-airflow turbulence as the perturbation prior.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Jul 8

AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models

arXiv:2607. 06485v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly deployed on infrared (IR) remote sensing imagery in security-critical settings, yet their adversarial robustness remains unexamined.

By Cong Su, Jiaju Han, Xuemeng Sun, Chengyin Hu, Qike Zhang, Jiujiang Guo, Yiwei Wei, Jiahuan Long
arXiv AI
Sep 3

InfraPatch: Cross-Task Targeted Grayscale Patch Attacks on Infrared-Adapted Vision-Language Models

InfraPatch is a white‑box, per‑instance framework that generates small grayscale patches to target infrared‑adapted vision‑language models (IR‑VLMs). The method optimizes a single‑channel patch within a 5% local‑area budget, using proxy‑guided placement and task‑adaptive objectives to induce desired behaviors in image classification, captioning, and binary visual question answering. Across ten IR‑VLM variants tested on synthetic infrared images, InfraPatch achieves targeted attack success rates ranging from 86% to 100%, revealing significant vulnerability differences among architectures and tasks.

By Chengyin Hu, Dingyi Lu, Jiaju Han, Xiang Chen, Weiwen Shi, Jiahuan Long, Yiwei Wei, Jiujiang Guo
arXiv AI
Aug 19

Training with synthetic data for drone detection in thermal imagery

The paper explores a synthetic-first training approach for detecting drones in medium- and long-wave infrared imagery, combining synthetic scene generation with fine-tuning on real data. It demonstrates that synthetic data can establish initial object representations, but real infrared data is crucial to close domain gaps and improve reliability. The study finds that aligning datasets has a greater impact on performance than increasing model size, and that semantic alignment in feature space is the strongest predictor of success, with radiometric factors like entropy and dynamic range also contributing.

By Tanel Liiv, Sander Soodla, Nzamba Bignoumba, Alma M. Liezenga, Toomas Pruuden
arXiv AI
Aug 20

Breaking the weakest link to evade vision language models

The paper investigates how Vision Language Models (VLMs) can be fooled by small, human‑imperceptible changes to images. It introduces a gradient‑based attack that targets only the vision encoder, reducing computational cost while still effectively disrupting both untargeted and targeted multimodal alignment. Experiments on open‑source VLMs such as Qwen2.5‑VL, Granite‑Vision, FastVLM, and Phi‑3.5‑Vision demonstrate that these perturbations can significantly alter the models’ textual outputs.

By Ilan Zini, Boussad Addad, Katarzyna Kapusta