arXiv AI By Suyoung Lee, Myungsub Choi

Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection Guardrails

Read the original on arXiv AI →

The paper introduces Mind2Web-Injection, a benchmark of 9,954 instruction‑screenshot pairs designed to evaluate whether vision‑language models (VLMs) use visual evidence when judging conflicts between on‑screen text and user instructions. Across six VLMs, models with similar overall precision can differ dramatically in Evidence‑Aligned Detection (EAD), revealing that verdict‑only metrics miss critical grounding failures. The authors also propose two training‑free interventions—ReadGate and CmdCompare—to improve evidence grounding and instruction‑command consistency, and argue for separate reporting of verdict correctness, evidence localization, and counterfactual responsiveness.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 21

Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs

arXiv:2607. 16311v1 Announce Type: cross Abstract: Vision-language models (VLMs) often answer visual questions using learned language and category priors rather than grounding their predictions in the image itself.

By Jingyu Sun, Jiachen Tu, Yuyang Xue, Yaoxin Jiang, Guoyi Xu, Zhengtao Yao, Rui Qian, Yizheng Sun, Hongpeng Zhou, Jingyuan Sun, Yan Lin
arXiv AI
Sep 11

AgentHijack: Visual Patch Attacks on Multimodal Computer-Use Agents

The paper introduces an end‑to‑end framework for testing image‑triggered command injection on computer‑use agents (CUAs). It demonstrates that a local visual patch can cause verifiable environmental effects through the entire pipeline—from screenshot input, through vision‑language‑model generation, action parsing, to execution—by training and deploying patches on GitHub Pages and a CSDN clone. Across five open‑source or publicly available GUI‑agent or VLM backends, the study reports 84.5% T‑ASR, 47.0% TAPR, and 20.3% E2E‑ASR success rates, with trajectory analysis revealing that some attacks first execute malicious terminal commands before continuing the original task.

By Zhihao Liu, Hongyu Sun, Zhiyuan Fu, Xiaonan Duan, Jice Wang, Shangru Zhao, Weizhi Meng, Wuxin Yang, Yangfan Zhou, Yuqing Zhang
arXiv AI
Aug 7

Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots

arXiv:2608. 05715v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly deployed as planners in robotic systems, where they translate natural-language commands into executable actions grounded in visual scene understanding.

By S. M . Bhagya P. Samarakoon, M. A. Viraj J. Muthugala, W. K. R. Sachinthana, Mohan Rajesh Elara