arXiv AI

VisualLeakBench: Reproducible Action-Boundary Propagation Failures in Vision-Language Agents

arXiv:2606. 07595v1 Announce Type: cross Abstract: Vision-language agents increasingly consume screenshots, documents, and user interfaces before writing to memory, sending messages, or invoking external tools.

arXiv AI
Sep 11

AgentHijack: Visual Patch Attacks on Multimodal Computer-Use Agents

The paper introduces an end‑to‑end framework for testing image‑triggered command injection on computer‑use agents (CUAs). It demonstrates that a local visual patch can cause verifiable environmental effects through the entire pipeline—from screenshot input, through vision‑language‑model generation, action parsing, to execution—by training and deploying patches on GitHub Pages and a CSDN clone. Across five open‑source or publicly available GUI‑agent or VLM backends, the study reports 84.5% T‑ASR, 47.0% TAPR, and 20.3% E2E‑ASR success rates, with trajectory analysis revealing that some attacks first execute malicious terminal commands before continuing the original task.

By Zhihao Liu, Hongyu Sun, Zhiyuan Fu, Xiaonan Duan, Jice Wang, Shangru Zhao, Weizhi Meng, Wuxin Yang, Yangfan Zhou, Yuqing Zhang
arXiv AI
Sep 7

When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs

The paper introduces MPS-Bench, a benchmark of 5,181 scenarios from 584 real-world images across 12 high-risk domains, each paired with a hidden user profile, to evaluate personalized safety in vision‑language models (VLMs). Eight leading VLMs were tested and found to almost always respond directly (86‑99%) without seeking missing context, scoring no higher than 2.6/5 on personalized safety. The authors identify a phenomenon called visual dominance, where visual information enters text representations early and suppresses textual risk signals, and propose PRISM, a lightweight input monitor that predicts when a query should be deferred, achieving 0.978 AUC and outperforming all tested models on the safety‑utility Pareto frontier.

By Edward Sun, Yuchen Wu, Zixian Ma, Eric Hanchen Jiang, Yijia Xiao, Xiaoyuan Yi, Ranjay Krishna, Wei Wang, Jindong Wang, Aylin Caliskan
arXiv AI
Aug 19

COMIC: Reference-Aware Safety Gating for Multimodal Large Language Models

The paper introduces COMIC, a reference‑aware safety gate designed for multimodal large language models (MLLMs). COMIC detects the operation requested by a user, identifies visual targets through OCR and open‑vocabulary proposals, and evaluates safety on explicit operation‑target pairs, using max‑risk aggregation and quality‑aware routing to decide whether to allow or block a request. Experiments on several open‑source MLLMs and jailbreak benchmarks show that COMIC improves robustness while maintaining benign utility and efficiency.

By Md Abdullahil Oaphy, Anhao Xiang, Zongxing Xie, Huayue Gu, Chenyu Wang, Honghui Xu
arXiv AI
Sep 7

Repeat-After-Me: Black-Box Adaptive Visual Prompt Injection

Repeat-After-Me is a black-box adaptive visual prompt injection technique that can reveal personally identifiable information or trigger malicious tool calls in both open-weight and commercial vision‑language models, achieving attack success rates above 80% on Qwen3.6‑27B and 47% on GPT‑5.5. The method works even when the benign user prompt is unrelated to the injected task and does not explicitly authorize it, and it retains significant effectiveness when transferred across models or optimized on surrogate systems. In a real‑world OpenClaw Discord deployment, a minimally injected image can overwrite TOOLS.md, enabling remote code execution and secret exfiltration.

By Sizhe Chen, Yu-Lin Tsai, Ivan Evtimov, Kamalika Chaudhuri, Raluca Ada Popa, David Wagner, Arman Zharmagambetov
arXiv AI
Sep 10

Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection Guardrails

The paper introduces Mind2Web-Injection, a benchmark of 9,954 instruction‑screenshot pairs designed to evaluate whether vision‑language models (VLMs) use visual evidence when judging conflicts between on‑screen text and user instructions. Across six VLMs, models with similar overall precision can differ dramatically in Evidence‑Aligned Detection (EAD), revealing that verdict‑only metrics miss critical grounding failures. The authors also propose two training‑free interventions—ReadGate and CmdCompare—to improve evidence grounding and instruction‑command consistency, and argue for separate reporting of verdict correctness, evidence localization, and counterfactual responsiveness.

By Suyoung Lee, Myungsub Choi