MobileWorldSafety is a benchmark that evaluates the safety of large language model–powered GUI agents on Android by exposing them to 142 real-world risk tasks involving environmental injection attacks. The benchmark uses a two‑stage verification pipeline—rule‑based checks for clear cases and an LLM judge for ambiguous ones—to distinguish safety failures from capability failures. Experiments on six agents show high vulnerability, with attack success rates between 40.4% and 66.9%, highlighting that current agents often fail to remain safe when faced with adversarial content presented as normal mobile context.
By Sujin Chen, Lijun Li, Tianyi Du, Jing Shao
arXiv:2601. 12349v3 Announce Type: replace-cross Abstract: Large multimodal model powered GUI agents are emerging as high-privilege operators on mobile platforms, entrusted to perceive screen content and inject inputs across application boundaries.
By Yi Qian, Kunwei Qian, Xingbang He, Ligeng Chen, Jikang Zhang, Tiantai Zhang, Haiyang Wei, Linzhang Wang, Hao Wu, Bing Mao
arXiv:2606. 12666v1 Announce Type: cross Abstract: Screenshot-based mobile GUI agents can operate ordinary smartphone apps through the same visual interface as a human user, but this capability also turns every screen observation into a privacy boundary.
By Siyu Shen, Fenghao Xu, Wenrui Diao, Kehuan Zhang
arXiv:2510. 24411v3 Announce Type: replace Abstract: Computer-using agents powered by Vision-Language Models (VLMs) have demonstrated human-like capabilities in operating digital environments like mobile platforms.
By Qiushi Sun, Mukai Li, Zhoumianze Liu, Zhihui Xie, Fangzhi Xu, Zhangyue Yin, Kanzhi Cheng, Zehao Li, Zichen Ding, Qi Liu, Zhiyong Wu, Zhuosheng Zhang, Ben Kao, Lingpeng Kong
AI agents are driving a new software paradigm, with the ability to autonomously call tools, extract information, manage memory, and complete tasks that span applications and data sources. Most existing end-user operating systems, however, are designed for application-centric workflows and offer little native support for AI agents.
AgentLens is a mobile GUI agent that adapts its visual communication during task execution, offering Full UI, Partial UI, and GenUI modalities. It uses a Virtual Display to allow background operation while selectively overlaying visual information. In a study with 21 participants, 85.7% preferred AgentLens, which also scored highest on usability and adoption intent.
By Jeonghyeon Kim, Byeongjun Joung, Junwon Lee, Joohyung Lee, Taehoon Min, Sunjae Lee
ADeptS-Bench is a new benchmark designed to assess the trustworthiness of Computer Use Agents (CUAs) across mobile and desktop devices. It consists of two streams: a Safety stream with paired benign and malicious tasks that embed visual threats, and a Disambiguation stream that tests whether agents seek clarification when instructions are ambiguous. Evaluation of seven models shows none consistently achieves high task success while keeping attack success low, and all models exhibit problematic behaviors such as unhesitant checkout on a $25K order and failure to detect a mislabeled factory reset button.
By Joy Chen, Alejandro Castillejo Munoz, Pierluca D'Oro, Yuxuan Sun, Chloe Evans, Joseph Tighe
arXiv:2609.22724v1 Announce Type: cross
Abstract: Mobile agents powered by foundation models now automate complex, multi-step workflows on real devices, but their trajectories can violate app-specifi...
By Changyue Jiang, Jiayi Wang, Xin Wen, Jiarun Dai, Geng Hong, Xudong Pan
arXiv:2609.23980v1 Announce Type: cross
Abstract: AI agents now report vulnerabilities faster than maintainers can review them. Reports often depend on security properties specific to the application...
By Andy K. Zhang, Ava Huang, Joey Ji, Wai Han, Thomas Qin, Nardos Demilew, Michael Tian-Yue Liu, Brian Song, Riya Dulepet, Brian Wang, Kyleen Liao, Cuiyuanxiu Chen, Nishka Kacheria, Andrew Wu, Pratham Rangwala, Xinjie Wang, Laura Gomezjurado Gonzalez, Anita Ding, Benjamin Yi, Daniel E. Ho, Dan Boneh, Dawn Song, Ion Stoica, Percy Liang
arXiv:2606. 20470v1 Announce Type: cross Abstract: Agentic AI systems increasingly rely on language-model components to interpret instructions, process external data, invoke tools, and coordinate with other agents.
By Reza Soosahabi, Vivek Namsani
arXiv:2605. 29486v2 Announce Type: replace-cross Abstract: A central bottleneck for phone-use agents is that controllable, reproducible environments covering real mobile behavior are hard to build at scale.
By Yuxuan Liu, Xin Lai, Junyi Li, Pengyuan Lyu, Jason, Yiduo Guo, Zhengyao Fang, Yang Ding, Yi Zhang, Weinong Wang, Huawen Shen, Xingran Zhou, Liang Wu, Fei Tang, Sunqi Fan, Shangpin Peng, Zheng Ruan, Anran Zhang, Chengquan Zhang, Han Hu, Benyou Wang, Ji-Rong Wen, Rui Yan, Zhengyang Tang
The paper introduces AnTrap, a benchmark that injects dynamic perturbations into Android GUI agent execution to evaluate robustness against runtime anomalies. It presents a taxonomy of anomalies across four layers—State, Thinking, Action, and Round—with ten subcategories, and a pipeline that maintains task solvability while adding realistic adversarial conditions. Experiments on 16 leading GUI models show universal vulnerability, and reinforcement learning can mitigate some traps but not deep contextual ones like state deadlock.
By Guo Gan, Yilun Zhao, Cong Chen, Jinbiao Wei, Tingyu Song, Zheyuan Yang, Lin Fu, Hong Zhou