MobileWorldSafety is a benchmark that evaluates the safety of large language model–powered GUI agents on Android by exposing them to 142 real-world risk tasks involving environmental injection attacks. The benchmark uses a two‑stage verification pipeline—rule‑based checks for clear cases and an LLM judge for ambiguous ones—to distinguish safety failures from capability failures. Experiments on six agents show high vulnerability, with attack success rates between 40.4% and 66.9%, highlighting that current agents often fail to remain safe when faced with adversarial content presented as normal mobile context.
By Sujin Chen, Lijun Li, Tianyi Du, Jing Shao
arXiv:2608. 07775v1 Announce Type: new Abstract: Mobile agents have achieved promising results on clean online benchmarks such as AndroidWorld, yet their performance often degrades sharply in real-world deployment due to environmental variations and imperfect interface conditions.
By Xiaoou Liu, Longchao Da, Hanyang Chen, Yuan Ling, Hua Wei
arXiv:2608.22847v1 Announce Type: new
Abstract: Vision-Language Models (VLMs) based GUI agents stand to benefit significantly from online reinforcement learning (RL). However, their training is bottl...
By Long Zhang, Yuhan Chen, Chaoran Zhang, Wanxia Cao, Kun Huang, Pengzhi Gao, Wei Liu, Jian Luan, Chenliang Li, Lixin Zou
Vision-Language Models (VLMs) based GUI agents stand to benefit significantly from online reinforcement learning (RL). However, their training is bottlenecked by two fundamental issues: current data s...
The paper introduces Replicant, a deep reinforcement learning framework that learns to evade malware detectors under a strict label‑only black‑box threat model. Replicant generates reusable policies for modifying malware samples and deciding when to query the target, and it transfers across different samples, detectors, and feature spaces. In experiments on seven Android malware detectors and three feature spaces, Replicant achieves a mean attack success rate of 78.8%, outperforming state‑of‑the‑art methods by 20.9%–39.2% and providing a stronger signal for adversarial training to harden detectors.
By Shae McFadden, Ilias Tsingenopoulos, Mario D'Onghia, Alexander Herzog, Myles Foley, Chris Hicks, Lorenzo Cavallaro, Fabio Pierazzi
arXiv:2605. 29486v2 Announce Type: replace-cross Abstract: A central bottleneck for phone-use agents is that controllable, reproducible environments covering real mobile behavior are hard to build at scale.
By Yuxuan Liu, Xin Lai, Junyi Li, Pengyuan Lyu, Jason, Yiduo Guo, Zhengyao Fang, Yang Ding, Yi Zhang, Weinong Wang, Huawen Shen, Xingran Zhou, Liang Wu, Fei Tang, Sunqi Fan, Shangpin Peng, Zheng Ruan, Anran Zhang, Chengquan Zhang, Han Hu, Benyou Wang, Ji-Rong Wen, Rui Yan, Zhengyang Tang