arXiv Machine Learning

Temporal Generalization and Explanation Stability of Control Flow Graph Neural Networks for Malware Detection

The paper evaluates how graph neural networks (GNNs) built on control flow graphs (CFGs) perform when trained on one time period and tested on a later one, using a strict temporal split. Twelve GNN variants and a flat-feature baseline were trained on 459 CFGs from 2024‑2025 and evaluated on 223 CFGs from 2026, revealing that the choice of message‑passing operator strongly affects robustness to distribution shift. Attribution stability varied by architecture, and the best-performing operator on the later corpus was also the hardest to explain, leading the authors to design a new architecture that matches its performance without search.

arXiv Machine Learning
Jul 28

EXE-Bench: Ranking the Tradeoffs of AI-based Windows Malware Detectors for Real-World Usability

arXiv:2607. 24177v1 Announce Type: cross Abstract: Due to the lack of systematic evaluations, we are not yet able to determine which AI-based Windows malware detector to deploy in production, since existing evaluations (i) differ in terms of data used for both training and testing; (ii) do not consider temporal analysis to showcase whether models withstand the passage of time; (iii) avoid security evaluations with adversarial attacks that could highlight their brittleness against content-injection attacks; and (iv) neglect the computational requirements for deployment, risking slow inference on endpoints.

By Andrea Ponte, Daniel Gibert, Matous Kozak, Dmitrijs Trizna, Maura Pintor, Battista Biggio, Fabio Roli, Luca Demetrio
arXiv Machine Learning
Sep 23

HYDRA: Proactive Android Malware Drift Adaptation via Hierarchical Graph Contrastive Learning

HYDRA is a proactive Android malware drift adaptation framework that learns drift‑invariant representations from hierarchically structured data. It combines fine‑grained Control Flow Graphs and coarse‑grained Function Call Graphs to model applications, then applies a cross‑domain contrastive learning objective to align historical and new data distributions. Experiments on large‑scale, time‑ordered malware datasets show HYDRA achieves lower false negative and false positive rates than state‑of‑the‑art baselines while needing up to 87.5% fewer labeled samples.

By Han Chen, Hanchen Wang, Hongmei Chen, Lu Qin, Wenjie Zhang, Ying Zhang
arXiv Machine Learning
Jun 10

Do Transformers Actually Help Intrusion Detection? A Temporal Sequence Evaluation on CIC-IDS2017

arXiv:2606. 11098v1 Announce Type: cross Abstract: Recent deep learning approaches for network intrusion detection increasingly incorporate temporal architectures such as recurrent networks and Transformers, often reporting near-perfect performance on CIC-IDS2017.

By Zach Moczkodan (Royal Military College of Canada, Kingston, Canada), Hany Ragab (Royal Military College of Canada, Kingston, Canada)