The paper introduces Motion Vision CAPTCHA (MVCAP), a new CAPTCHA framework that relies on motion-defined foreground structures to create challenges that are only solvable through temporal analysis of a dynamic background. MVCAP is implemented in three progressive levels—coherent motion, structural motion, and biological motion—and evaluated using the MVCAP-Bench, a browser-based benchmark with 600 live CAPTCHA instances. Human participants achieve 99.6% accuracy, whereas the best GUI agent scores only 16.8%, highlighting a significant human–agent perception gap and demonstrating that dynamic background camouflage is the key difficulty.
By Zeyu Zhang, Dingyi Rong, Zijian Chen, Zicheng Zhang, Xiongkuo Min, Guangtao Zhai
arXiv:2603. 23559v2 Announce Type: replace-cross Abstract: GUI agents are rapidly shifting from multi-module pipelines to end-to-end, native vision-language models (VLMs) that perceive raw screenshots and directly interact with digital devices.
By Yuxi Chen, Haoyu Zhai, Chenkai Wang, Rui Yang, Lingming Zhang, Gang Wang, Huan Zhang
arXiv:2512. 02318v4 Announce Type: replace-cross Abstract: This paper studies how multimodal large language models (MLLMs) undermine the security guarantees of visual CAPTCHA.
By Junyu Wang, Changjia Zhu, Yuanbo Zhou, Lingyao Li, Xu He, Mingkui Wei, Junjie Xiong
arXiv:2607. 18659v1 Announce Type: cross Abstract: LLM-based browser agents are rapidly changing the threat landscape for web security.
By Behzad Ousat, Nikita Turkmen, Lalchandra Rampersaud, Dillan Bailey, Amin Kharraz
LLM-based browser agents are rapidly changing the threat landscape for web security. Unlike traditional automation frameworks that execute predefined scripts, these agents can autonomously navigate websites, reason about page content, and interact with web interfaces using natural-language instructions.
arXiv:2605.06524v3 Announce Type: replace
Abstract: Reliable human-machine discrimination is becoming increasingly important as Large Language Models and autonomous agents are deployed in online sett...
By Milena Rmus, Mathew D. Hardy, Thomas L. Griffiths, Mayank Agrawal
arXiv:2608. 14870v1 Announce Type: new Abstract: Long-horizon agents can execute continuously, but human attention remains intermittent and scarce.
By Chen Chen, Zhehuai Chen
arXiv:2607. 14443v1 Announce Type: new Abstract: Computer-use agents are becoming capable software operators, but their interface to desktop applications is still often a brittle motor layer: they look at screenshots, predict coordinates, click, and hope that the visible state changed as intended.
By Yong Liu, Zhenyi Zhong, Zhanpeng Shi
arXiv:2607. 10878v1 Announce Type: new Abstract: AI agents are evolving from answer engines into persistent teams that use tools, delegate work, learn from experience, and modify the artifacts that shape their future behavior.
By Yuma Ichikawa, Yamato Arai, Kosaku Kimura, Akira Sakai, Hiromichi Kobashi
arXiv:2606. 07594v1 Announce Type: new Abstract: Personal AI agents must increasingly operate across APIs, shells, web surfaces, and desktop GUIs, yet many systems remain tuned to a single interface and offer limited support for user teaching and auditability.
By Bo Zhang, Borui Zhang, Chenghao Jiang, Minglei Shi, Xiaofeng Wang, Zheng Zhu, Jie Zhou, Jiwen Lu
arXiv:2608. 05715v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly deployed as planners in robotic systems, where they translate natural-language commands into executable actions grounded in visual scene understanding.
By S. M . Bhagya P. Samarakoon, M. A. Viraj J. Muthugala, W. K. R. Sachinthana, Mohan Rajesh Elara
The paper surveys Multimodal Code Intelligence, focusing on tasks where code is generated, edited, refined, or reasoned about under visually grounded inputs such as screenshots, charts, and videos. It categorizes the field by the role of code—rendered artifact, editable structure, intermediate reasoning trace, or executable tool interface—and organizes benchmarks into four domains: Graphical User Interface, Scientific Visualization, Structured Graphics, and Frontier Tasks and Frameworks. The authors argue that reliable evaluation must include evidence of semantics and interaction beyond visual fidelity, and propose four verification-centered research directions to advance the field toward evidence-grounded executable systems.
By Xuanle Zhao, Qiushi Sun, Jingyu Xiao, Xuexin Liu, Haoyue Yang, Qiaosheng Chen, Xianzhen Luo, Jing Huang, Yufeng Zhong, Lei Chen, Shuai Fu, Zhenlin Wei, Jinhe Bi, Lei Jiang, Haibo Qiu, Siqi Yang, Peng Shi, Jian Hu, Zhixiong Zeng