arXiv:2601. 12349v3 Announce Type: replace-cross Abstract: Large multimodal model powered GUI agents are emerging as high-privilege operators on mobile platforms, entrusted to perceive screen content and inject inputs across application boundaries.
By Yi Qian, Kunwei Qian, Xingbang He, Ligeng Chen, Jikang Zhang, Tiantai Zhang, Haiyang Wei, Linzhang Wang, Hao Wu, Bing Mao
arXiv:2512. 20872v2 Announce Type: replace-cross Abstract: Function call graphs (FCGs) have emerged as a powerful abstraction for malware detection, capturing the behavioral structure of applications beyond surface-level signatures.
By Jakir Hossain, Jue Guo, Gurvinder Singh, Lukasz Ziarek, Ahmet Erdem Sar{\i}y\"uce
arXiv:2609.23980v1 Announce Type: cross
Abstract: AI agents now report vulnerabilities faster than maintainers can review them. Reports often depend on security properties specific to the application...
By Andy K. Zhang, Ava Huang, Joey Ji, Wai Han, Thomas Qin, Nardos Demilew, Michael Tian-Yue Liu, Brian Song, Riya Dulepet, Brian Wang, Kyleen Liao, Cuiyuanxiu Chen, Nishka Kacheria, Andrew Wu, Pratham Rangwala, Xinjie Wang, Laura Gomezjurado Gonzalez, Anita Ding, Benjamin Yi, Daniel E. Ho, Dan Boneh, Dawn Song, Ion Stoica, Percy Liang
arXiv:2605. 09028v3 Announce Type: replace Abstract: Machine learning-based Android malware detectors often fail in real-world deployment due to domain shift, where models trained on one data source perform poorly on applications from another.
By Md Rafid Islam
arXiv:2503.11841v2 Announce Type: replace-cross
Abstract: Machine Learning (ML) malware detectors rely heavily on crowd-sourced AntiVirus (AV) labels, with platforms like VirusTotal serving as truste...
By Tianwei Lan, Luca Demetrio, Farid Nait-Abdesselam, Yufei Han, Simone Aonzo
MobileWorldSafety is a benchmark that evaluates the safety of large language model–powered GUI agents on Android by exposing them to 142 real-world risk tasks involving environmental injection attacks. The benchmark uses a two‑stage verification pipeline—rule‑based checks for clear cases and an LLM judge for ambiguous ones—to distinguish safety failures from capability failures. Experiments on six agents show high vulnerability, with attack success rates between 40.4% and 66.9%, highlighting that current agents often fail to remain safe when faced with adversarial content presented as normal mobile context.
By Sujin Chen, Lijun Li, Tianyi Du, Jing Shao
arXiv:2608. 03250v1 Announce Type: cross Abstract: The rapid advancement of modern technology has led to a significant increase in the use of smart devices, such as smartphones and tablets, resulting in the widespread adoption of mobile applications.
By Md Faisal Ahmed, Zarin Tasnim Biash, Abu Raihan Shakil, Ahmed Ann Noor Ryen, Arman Hossain, Faisal Bin Ashraf, Muhammad Iqbal Hossain
The paper introduces Replicant, a deep reinforcement learning framework that learns to evade malware detectors under a strict label‑only black‑box threat model. Replicant generates reusable policies for modifying malware samples and deciding when to query the target, and it transfers across different samples, detectors, and feature spaces. In experiments on seven Android malware detectors and three feature spaces, Replicant achieves a mean attack success rate of 78.8%, outperforming state‑of‑the‑art methods by 20.9%–39.2% and providing a stronger signal for adversarial training to harden detectors.
By Shae McFadden, Ilias Tsingenopoulos, Mario D'Onghia, Alexander Herzog, Myles Foley, Chris Hicks, Lorenzo Cavallaro, Fabio Pierazzi
arXiv:2508. 10031v2 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) have shown significant advancements in performance, various jailbreak attacks have posed growing safety and ethical risks.
By Jinhwa Kim, Ian G. Harris
The paper introduces a new benchmark for assessing out-of-distribution robustness in graph-based Android malware classifiers, highlighting that current models drop up to 45% accuracy on unseen malware variants. It presents two scenarios—MalNet-Tiny-Common for covariate shift and MalNet-Tiny-Distinct for domain shift—and identifies a limitation in existing benchmarks that rely solely on structure-only function call graphs. To address this, the authors propose a semantic enrichment framework that augments graph topology with function-level attributes and LLM-based code embeddings, demonstrating that this data-centric approach improves robustness under distribution shift and complements model-based methods.
By Ngoc N. Tran, Anwar Said, Waseem Abbas, Tyler Derr, Xenofon D. Koutsoukos
arXiv:2604. 23025v2 Announce Type: replace-cross Abstract: Android malware detectors built with machine learning often suffer from temporal bias: models are trained and evaluated without respecting apps' actual release times, inflating accuracy and weakening real-world robustness.
By Annan Fu, Hao Pei, Maryam Tanha
arXiv:2507. 18313v2 Announce Type: replace Abstract: Malware evolves rapidly, forcing machine learning-based detectors to be continuously updated.
By Daniele Ghiani, Daniele Angioni, Giorgio Piras, Angelo Sotgiu, Luca Minnei, Srishti Gupta, Maura Pintor, Fabio Roli, Battista Biggio