arXiv AI

COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers

arXiv:2512. 02318v4 Announce Type: replace-cross Abstract: This paper studies how multimodal large language models (MLLMs) undermine the security guarantees of visual CAPTCHA.

arXiv AI
Jun 30

CAPTCHA Solving for Native GUI Agents: Automated Reasoning-Action Data Generation and Self-Corrective Training

arXiv:2603. 23559v2 Announce Type: replace-cross Abstract: GUI agents are rapidly shifting from multi-module pipelines to end-to-end, native vision-language models (VLMs) that perceive raw screenshots and directly interact with digital devices.

By Yuxi Chen, Haoyu Zhai, Chenkai Wang, Rui Yang, Lingming Zhang, Gang Wang, Huan Zhang
arXiv AI
Jun 2

HLL: Can Agents Cross Humanity's Last Line of Verification?

arXiv:2606. 02449v1 Announce Type: new Abstract: Multimodal agents are increasingly expected to operate interfaces on behalf of users, raising a central deployment question: can they truly substitute for humans in workflows that services deliberately protect against automation?

By Xinhao Song, Su Su, Sirui Song, Hongliang Wu, Wen Shen, Zhihua Wei, Gongshen Liu, Linfeng Zhang, Dongrui Liu
arXiv Computer Vision
Sep 24

Invisible in Space, Visible in Time: Motion Vision CAPTCHA against GUI Agents

The paper introduces Motion Vision CAPTCHA (MVCAP), a new CAPTCHA framework that relies on motion-defined foreground structures to create challenges that are only solvable through temporal analysis of a dynamic background. MVCAP is implemented in three progressive levels—coherent motion, structural motion, and biological motion—and evaluated using the MVCAP-Bench, a browser-based benchmark with 600 live CAPTCHA instances. Human participants achieve 99.6% accuracy, whereas the best GUI agent scores only 16.8%, highlighting a significant human–agent perception gap and demonstrating that dynamic background camouflage is the key difficulty.

By Zeyu Zhang, Dingyi Rong, Zijian Chen, Zicheng Zhang, Xiongkuo Min, Guangtao Zhai
arXiv AI
Sep 12

Deep-Fake CAPTCHA: Mitigating Next-Generation Social Engineering Attacks

The paper introduces DF‑CAPTCHA, an active defense that asks callers to complete simple challenge‑response tasks during voice and video calls. By evaluating responses on realism, identity consistency, task completion, and response time, the system can detect real‑time deepfake impersonations. Experiments across audio and video modalities show that DF‑CAPTCHA outperforms passive artifact‑search methods, achieving high accuracy and demonstrating the effectiveness of challenge‑based verification against next‑generation social engineering attacks.

By Guy Frankovits, Lior Yasur, Fred M. Grabovski, Yisroel Mirsky
arXiv Machine Learning
Sep 21

OverThink: Slowdown Attacks on Reasoning LLMs

The paper introduces OverThink, a slowdown attack that forces reasoning language models (RLMs) to produce many more reasoning tokens while still giving correct answers. By injecting decoy reasoning problems—such as Markov decision processes, language translation, or graphic comprehension—into the model’s context, attackers can dramatically increase token generation (up to 46× on SQuAD and 17× on coding agents). The study evaluates the attack on both proprietary and open-source RLMs across multiple datasets, explores multimodal and coding‑agent variants, and tests several defenses, concluding that defending against OverThink is challenging and that newer RLMs are even more vulnerable due to higher per‑token costs and increased reasoning token usage.

By Abhinav Kumar, Jaechul Roh, Ali Naseh, Marzena Karpinska, Mohit Iyyer, Amir Houmansadr, Eugene Bagdasarian
Hugging Face Trending Papers
Aug 19

Breaking the weakest link to evade vision language models

The paper investigates how Vision Language Models (VLMs) can be fooled by tiny, human‑imperceptible changes to images. It introduces a gradient‑based attack that targets only the vision encoder, reducing computational cost while still effectively disrupting both untargeted and targeted multimodal interpretations. Experiments on open‑source VLMs such as Qwen2.5‑VL, Granite‑Vision, FastVLM, and Phi‑3.5‑Vision demonstrate that these small perturbations can dramatically alter the models’ textual outputs.

arXiv AI
Aug 7

Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots

arXiv:2608. 05715v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly deployed as planners in robotic systems, where they translate natural-language commands into executable actions grounded in visual scene understanding.

By S. M . Bhagya P. Samarakoon, M. A. Viraj J. Muthugala, W. K. R. Sachinthana, Mohan Rajesh Elara