arXiv Machine Learning

Do Agents Dream of False Memories? Black-box Visual Attacks on Long-term Memory in Multimodal AI Agents

arXiv:2607. 15657v1 Announce Type: cross Abstract: Multimodal AI agents increasingly rely on persistent long-term memory to ground generation in past visual and textual episodes.

arXiv AI
4d ago

Render Before Reading: Visual Rendering as a Prompt Injection Defense

The paper investigates how multimodal large language models are more susceptible to prompt injection when adversarial instructions are presented as text rather than as non-textual inputs like images. It proposes a training‑free defense that renders untrusted payloads into typographic images (or audio) before they reach the model, a method called Pictionary. Experiments on ten models and two benchmarks show that this approach significantly lowers attack success rates while maintaining normal functionality, and that fine‑tuning on image‑rendered instructions can further reduce the modality gap.

By Jie Zhang, Andrei Baroian, Jan N. van Rijn, Avital Shafran, Florian Tram\`{e}r
arXiv AI
Aug 19

Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs

The paper introduces MemJack, a memory‑augmented multi‑agent framework that automatically generates jailbreak attacks on Vision‑Language Models (VLMs) using benign natural images as visual anchors. MemJack discovers visual anchors, camouflages them semantically, evaluates responses, repairs via reflection, and replans dynamically, forming a closed‑loop attack pipeline. The authors also create MemJack‑Bench, a dataset of over 113,000 interactive multimodal jailbreak trajectories, and show that MemJack achieves a 71.48% attack success rate against Qwen3‑VL‑Plus, reaching 90% under extended budgets, outperforming other baselines on natural‑image evaluation.

By Jianhao Chen, Haoyang Chen, Hanjie Zhao, Haozhe Liang, Zheng Wang, Tieyun Qian
arXiv AI
Aug 19

DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization

DiSCO is a zero‑shot, black‑box defense for text‑to‑image models that operates solely at the prompt level. It expands prompts with a distribution‑guided suffix using beam search and contrastive scoring against safe and unsafe image pools generated by the target model, iteratively refining until safe content is produced. The method improves safety on the I2P benchmark under various red‑teaming attacks, reducing attack success rates by 37.7% and 25.13% while preserving semantic fidelity and image coherence.

By Tong Zhang, Motasem Alfarra, Carlos Hinojosa, Christos Louizos, Bernard Ghanem
arXiv AI
Aug 24

Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation

Vis-Poison is a novel attack that poisons multimodal retrieval-augmented generation systems by inserting attacker-controlled images as visual evidence, without altering any textual metadata. The attack uses an automated multi-agent approach to create visually plausible poisoned images and has been tested on two multimodal RAG pipelines, four embedding models, and six generation models. In black-box settings, Vis-Poison achieves an end-to-end success rate between 40.16% and 65.40% against 30,000-entry knowledge bases, and remains effective against various multimodal large language models with an average success rate above 60%.

By Rujin Liang, Zhongpu Chen, Yuhao Lei, Xin Miao
arXiv AI
Aug 20

Breaking the weakest link to evade vision language models

The paper investigates how Vision Language Models (VLMs) can be fooled by small, human‑imperceptible changes to images. It introduces a gradient‑based attack that targets only the vision encoder, reducing computational cost while still effectively disrupting both untargeted and targeted multimodal alignment. Experiments on open‑source VLMs such as Qwen2.5‑VL, Granite‑Vision, FastVLM, and Phi‑3.5‑Vision demonstrate that these perturbations can significantly alter the models’ textual outputs.

By Ilan Zini, Boussad Addad, Katarzyna Kapusta
arXiv AI
Sep 10

Discovering Natural Transformation Vulnerabilities in Black-Box Vision Models

The paper introduces Adversarial Scenario Attack (ASA), a query‑based black‑box method that discovers natural transformation vulnerabilities in vision models by exploring background, weather, and material/color edits via a multimodal language model and a text‑guided generative editor. ASA outperforms previous query‑based generative attacks on ImageNet classifiers, achieving higher success rates with fewer queries while maintaining perceptual quality. The approach also shows image‑level and prompt‑level transferability, indicating reusable vulnerabilities across models and images.

By Dongsu Song, DaeYun GO, Jay Hoon Jung
Hugging Face Trending Papers
Aug 19

Breaking the weakest link to evade vision language models

The paper investigates how Vision Language Models (VLMs) can be fooled by tiny, human‑imperceptible changes to images. It introduces a gradient‑based attack that targets only the vision encoder, reducing computational cost while still effectively disrupting both untargeted and targeted multimodal interpretations. Experiments on open‑source VLMs such as Qwen2.5‑VL, Granite‑Vision, FastVLM, and Phi‑3.5‑Vision demonstrate that these small perturbations can dramatically alter the models’ textual outputs.

arXiv AI
Jul 21

Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models

arXiv:2607. 17779v1 Announce Type: new Abstract: Text-to-Image (T2I) generative models have achieved remarkable progress in synthesizing high-quality visual content, yet they remain vulnerable to adversarial misuse, particularly in generating Not-Safe-For-Work (NSFW) images.

By Dongdong Yang, Deyue Zhang, Zhao Liu, Zonghao Ying, Wenzhuo Xu, Jiankai Jin, Xiangzheng Zhang, Quanchen Zou