arXiv AI By Bojing Li, Duo Zhong, Prajna Bhandary, Raguvir S, Charles Maxa, Robert J Joyce, Charles Nicholas

MASCOT-Android: A Curated Dataset and Automated Collection Pipeline for Android Malware Source Code Specimens

Read the original on arXiv AI →

arXiv:2606. 16072v1 Announce Type: cross Abstract: Compared with binaries and decompiled code, malware source code more directly reflects the attackers' original intent.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 18

Evaluating Out-of-Distribution Robustness in Graph-Based Android Malware Classification: A New Principled Benchmark

The paper introduces a new benchmark for assessing out-of-distribution robustness in graph-based Android malware classifiers, highlighting that current models drop up to 45% accuracy on unseen malware variants. It presents two scenarios—MalNet-Tiny-Common for covariate shift and MalNet-Tiny-Distinct for domain shift—and identifies a limitation in existing benchmarks that rely solely on structure-only function call graphs. To address this, the authors propose a semantic enrichment framework that augments graph topology with function-level attributes and LLM-based code embeddings, demonstrating that this data-centric approach improves robustness under distribution shift and complements model-based methods.

By Ngoc N. Tran, Anwar Said, Waseem Abbas, Tyler Derr, Xenofon D. Koutsoukos
arXiv AI
Jun 3

Large Byte Model: Teaching Language Models About Compiled Code

arXiv:2606. 02834v1 Announce Type: cross Abstract: Malware analysis starts with the raw bytes of an executable program, and tools to "lift" these to higher-level representations, such as assembly, are expensive and subject to error.

By Florian St\"ortz, Catalin-Andrei Stan, Alexandru Dinu, Sandra Servia-Rodr\'iguez, Mihaela Gaman, Calin Miron, Edward Raff