Repeat-After-Me is a black-box adaptive visual prompt injection technique that can reveal personally identifiable information or trigger malicious tool calls in both open-weight and commercial vision‑language models, achieving attack success rates above 80% on Qwen3.6‑27B and 47% on GPT‑5.5. The method works even when the benign user prompt is unrelated to the injected task and does not explicitly authorize it, and it retains significant effectiveness when transferred across models or optimized on surrogate systems. In a real‑world OpenClaw Discord deployment, a minimally injected image can overwrite TOOLS.md, enabling remote code execution and secret exfiltration.
By Sizhe Chen, Yu-Lin Tsai, Ivan Evtimov, Kamalika Chaudhuri, Raluca Ada Popa, David Wagner, Arman Zharmagambetov
arXiv:2606. 07595v1 Announce Type: cross Abstract: Vision-language agents increasingly consume screenshots, documents, and user interfaces before writing to memory, sending messages, or invoking external tools.
By Youting Wang, Yuan Tang, Yitian Qian, Chen Zhao
arXiv:2508. 08521v2 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) are increasingly being used in a broad range of applications, bringing their security and behavioral control to the forefront.
By Mansi Phute, Ravikumar Balakrishnan
arXiv:2604. 08005v2 Announce Type: replace Abstract: Advancements in multimodal foundation models have enabled the development of Computer Use Agents (CUAs) capable of autonomously interacting with GUI environments.
By Dominik Seip, Matthias Hein
Computer use agents (CUAs) are vision-language models that perceive a screen and act on a real operating system through mouse, keyboard, and terminal, and they are increasingly deployed to automate ev...
arXiv:2608.30207v1 Announce Type: cross
Abstract: Computer use agents (CUAs) are vision-language models that perceive a screen and act on a real operating system through mouse, keyboard, and terminal...
By Chen Xiong, Zhiyuan He, Pin-Yu Chen, Stjepan Picek, Tsung-Yi Ho
Recent advancements in Image-to-Video (I2V) generation have transformed input images from simple appearance references into interactive control interfaces where visual cues such as arrows, sketches, and emojis orchestrate complex video dynamics with unprecedented controllability. However, these seemingly innocuous static cues can be interpreted by models as executable temporal instructions, unfolding into harmful actions in the generated videos.
The paper introduces Mind2Web-Injection, a benchmark of 9,954 instruction‑screenshot pairs designed to evaluate whether vision‑language models (VLMs) use visual evidence when judging conflicts between on‑screen text and user instructions. Across six VLMs, models with similar overall precision can differ dramatically in Evidence‑Aligned Detection (EAD), revealing that verdict‑only metrics miss critical grounding failures. The authors also propose two training‑free interventions—ReadGate and CmdCompare—to improve evidence grounding and instruction‑command consistency, and argue for separate reporting of verdict correctness, evidence localization, and counterfactual responsiveness.
By Suyoung Lee, Myungsub Choi
The paper introduces MemJack, a memory‑augmented multi‑agent framework that automatically generates jailbreak attacks on Vision‑Language Models (VLMs) using benign natural images as visual anchors. MemJack discovers visual anchors, camouflages them semantically, evaluates responses, repairs via reflection, and replans dynamically, forming a closed‑loop attack pipeline. The authors also create MemJack‑Bench, a dataset of over 113,000 interactive multimodal jailbreak trajectories, and show that MemJack achieves a 71.48% attack success rate against Qwen3‑VL‑Plus, reaching 90% under extended budgets, outperforming other baselines on natural‑image evaluation.
By Jianhao Chen, Haoyang Chen, Hanjie Zhao, Haozhe Liang, Zheng Wang, Tieyun Qian
arXiv:2509.11250v3 Announce Type: replace-cross
Abstract: Graphical User Interface (GUI) agents are increasingly deployed to interact with online web services, yet their exposure to open-world conten...
By Yitong Zhang, Ximo Li, Liyi Cai, Jia Li
arXiv:2609. 09404v1 Announce Type: cross Abstract: Agentic AI frameworks let a language model plan, keep memory, and call tools that reach real files, mail, and services.
By Viet K. Nguyen, Mohammad I. Husain
arXiv:2606.23189v2 Announce Type: replace-cross
Abstract: Computer-use agents (CUAs) now act on a user's behalf across personal applications such as email, calendars, and to-do lists. This cross-appl...
By Anmol Goel, Iryna Gurevych