HYDRA is a proactive Android malware drift adaptation framework that learns drift‑invariant representations from hierarchically structured data. It combines fine‑grained Control Flow Graphs and coarse‑grained Function Call Graphs to model applications, then applies a cross‑domain contrastive learning objective to align historical and new data distributions. Experiments on large‑scale, time‑ordered malware datasets show HYDRA achieves lower false negative and false positive rates than state‑of‑the‑art baselines while needing up to 87.5% fewer labeled samples.
By Han Chen, Hanchen Wang, Hongmei Chen, Lu Qin, Wenjie Zhang, Ying Zhang
arXiv:2608. 13465v1 Announce Type: cross Abstract: Concept drift refers to changes over time in the statistical properties of data, as compared to the data that was used to train a learning model.
By Christofer Washington Berruz Chungata, Martin Jurecek, Katerina Potika, William B. Andreopoulos, Mark Stamp
The paper addresses the challenge of adapting malware detection systems to new threats without retraining from scratch, focusing on the Few-Shot Class-Incremental Learning (FSCIL) setting. It proposes a hybrid framework that uses a self-supervised learning backbone pre-trained on malware packets, incorporates Low-Rank Adaptation (LoRA) to adapt the model while preserving core representations, and employs a prototype-based classification head for incremental sessions. Experiments on multiple datasets show that this approach consistently outperforms existing FSCIL baselines and achieves state-of-the-art performance.
By Kyle Stein, Guillermo Francia, III Eman El-Sheikh, Andrew Arash Mahyari
arXiv:2507. 18313v2 Announce Type: replace Abstract: Malware evolves rapidly, forcing machine learning-based detectors to be continuously updated.
By Daniele Ghiani, Daniele Angioni, Giorgio Piras, Angelo Sotgiu, Luca Minnei, Srishti Gupta, Maura Pintor, Fabio Roli, Battista Biggio
The paper introduces a new benchmark for assessing out-of-distribution robustness in graph-based Android malware classifiers, highlighting that current models drop up to 45% accuracy on unseen malware variants. It presents two scenarios—MalNet-Tiny-Common for covariate shift and MalNet-Tiny-Distinct for domain shift—and identifies a limitation in existing benchmarks that rely solely on structure-only function call graphs. To address this, the authors propose a semantic enrichment framework that augments graph topology with function-level attributes and LLM-based code embeddings, demonstrating that this data-centric approach improves robustness under distribution shift and complements model-based methods.
By Ngoc N. Tran, Anwar Said, Waseem Abbas, Tyler Derr, Xenofon D. Koutsoukos
arXiv:2503.11841v2 Announce Type: replace-cross
Abstract: Machine Learning (ML) malware detectors rely heavily on crowd-sourced AntiVirus (AV) labels, with platforms like VirusTotal serving as truste...
By Tianwei Lan, Luca Demetrio, Farid Nait-Abdesselam, Yufei Han, Simone Aonzo