arXiv AI

Concept Drift Detection and Adaptive Retraining of Malware Classification Models

arXiv:2608. 13465v1 Announce Type: cross Abstract: Concept drift refers to changes over time in the statistical properties of data, as compared to the data that was used to train a learning model.

arXiv Machine Learning
Sep 23

HYDRA: Proactive Android Malware Drift Adaptation via Hierarchical Graph Contrastive Learning

HYDRA is a proactive Android malware drift adaptation framework that learns drift‑invariant representations from hierarchically structured data. It combines fine‑grained Control Flow Graphs and coarse‑grained Function Call Graphs to model applications, then applies a cross‑domain contrastive learning objective to align historical and new data distributions. Experiments on large‑scale, time‑ordered malware datasets show HYDRA achieves lower false negative and false positive rates than state‑of‑the‑art baselines while needing up to 87.5% fewer labeled samples.

By Han Chen, Hanchen Wang, Hongmei Chen, Lu Qin, Wenjie Zhang, Ying Zhang
arXiv Machine Learning
Sep 24

Enhancing Multiclass Malware Classification in Resource-Constrained Environments

The paper presents a lightweight machine‑learning approach for multi‑class malware detection on resource‑constrained devices. Using a LightGBM classifier with SMOTE oversampling, SOM‑US undersampling, and Genetic‑Algorithm feature selection, the authors achieve 89.1 % accuracy on four malware families and 76 % on 16 individual malware types. A second Random‑Forest model further improves family classification to 91.2 % and individual classification to 78.7 %.

By Abdul Khalek Alve, Alif Rahman, Saadman Zaman, Sazzad Hossen Himel, Muhammad Iqbal Hossain
arXiv AI
Aug 25

Adapter-Based Few-Shot Continual Learning for Malicious Packet Recognition

The paper addresses the challenge of adapting malware detection systems to new threats without retraining from scratch, focusing on the Few-Shot Class-Incremental Learning (FSCIL) setting. It proposes a hybrid framework that uses a self-supervised learning backbone pre-trained on malware packets, incorporates Low-Rank Adaptation (LoRA) to adapt the model while preserving core representations, and employs a prototype-based classification head for incremental sessions. Experiments on multiple datasets show that this approach consistently outperforms existing FSCIL baselines and achieves state-of-the-art performance.

By Kyle Stein, Guillermo Francia, III Eman El-Sheikh, Andrew Arash Mahyari