arXiv:2607. 24177v1 Announce Type: cross Abstract: Due to the lack of systematic evaluations, we are not yet able to determine which AI-based Windows malware detector to deploy in production, since existing evaluations (i) differ in terms of data used for both training and testing; (ii) do not consider temporal analysis to showcase whether models withstand the passage of time; (iii) avoid security evaluations with adversarial attacks that could highlight their brittleness against content-injection attacks; and (iv) neglect the computational requirements for deployment, risking slow inference on endpoints.
By Andrea Ponte, Daniel Gibert, Matous Kozak, Dmitrijs Trizna, Maura Pintor, Battista Biggio, Fabio Roli, Luca Demetrio
arXiv:2606. 30819v1 Announce Type: cross Abstract: Generative AI has emerged as a significant cybersecurity threat, with several recent attack campaigns leveraging LLMs to generate code for malicious purposes via scripting languages such as PowerShell.
By Luciano Pianese, Vittorio Orbinato, Pietro Liguori, Roberto Natella
arXiv:2606. 20436v1 Announce Type: cross Abstract: Malware analysts often inspect compiled binaries through decompiled pseudo-C, when source code is unavailable.
By Bercan Turkmen, Vyas Raina
arXiv:2606. 02834v1 Announce Type: cross Abstract: Malware analysis starts with the raw bytes of an executable program, and tools to "lift" these to higher-level representations, such as assembly, are expensive and subject to error.
By Florian St\"ortz, Catalin-Andrei Stan, Alexandru Dinu, Sandra Servia-Rodr\'iguez, Mihaela Gaman, Calin Miron, Edward Raff
arXiv:2606. 30572v1 Announce Type: cross Abstract: Malware classification remains a challenging problem due to its inherent heterogeneity, the presence of packed binaries, and the diverse distribution of malware families.
By Jithin S., Roshin Sleeba C., Anvin Mariya P. B., Asmitha K. A., Vinod P., Serena Nicolazzo, Antonino Nocera
arXiv:2509. 14335v2 Announce Type: replace-cross Abstract: Automated malware classifiers achieve strong detection performance, but auditing requires more than flagging a sample: analysts must explain malicious behaviors and justify them with code evidence.
By Xinran Zheng, Xingzhi Qian, Yiling He, Shuo Yang, Lorenzo Cavallaro
SCRIPTIOC-BENCH is a benchmark designed to evaluate how well large language models can statically extract indicators of compromise (IOCs) from script-based malware. It contains 634 manually verified JavaScript, PowerShell, and VBScript samples and covers four IOC types—URLs, domains, IP addresses, and filesystem artifacts—while distinguishing between directly exposed and encoded indicators. Experiments show that even the best models achieve only 65.4 F1, and a false‑positive taxonomy is introduced to analyze error patterns, with two mitigations (deterministic string utilities and task‑specific adaptation) improving precision and shifting errors toward sample‑grounded mismatches.
By Hanna Kim, Jian Cui, Minkyoo Song, Hwanjo Heo, Seungwon Shin, Kimin Lee, Xiaojing Liao
arXiv:2609.14593v1 Announce Type: cross
Abstract: Living-Off-the-Land (LOTL) is the dominant evasion technique of Advanced Persistent Threat (APT) actors, exploiting legitimate Windows utilities to c...
By Ahad Bin Islam Shoeb, Kamrul Hasan, Jamal Uddin Tanvin, Liang Hong, Imtiaz Ahmed, Md Arif Billah, Al Amin
arXiv:2608. 02671v1 Announce Type: cross Abstract: Malware detection using Hardware Performance Counters (HPC) has emerged as a promising solution to improve the security of computing systems as a complement to antivirus software.
By Alireza Abolhasani Zeraatkar, Parnian Shabani Kamran, Inderpreet Kaur, Nagabindu Ramu, Tyler Sheaves, Hussain Al-Asaad
arXiv:2610.01893v1 Announce Type: cross
Abstract: By 2030, Internet of Things (IoT) devices are projected to reach 40 billion, with fast-paced technological advancements in fields such as industry, h...
By Emmanuela Andam, Rana Shaaban, Emanuel Grant, Naima Kaabouch
Script-based malware remains a prevalent attack technique. These scripts often contain indicators of compromise (IOCs) that provide actionable threat intelligence. However, statically recovering such...
The paper introduces a new benchmark for assessing out-of-distribution robustness in graph-based Android malware classifiers, highlighting that current models drop up to 45% accuracy on unseen malware variants. It presents two scenarios—MalNet-Tiny-Common for covariate shift and MalNet-Tiny-Distinct for domain shift—and identifies a limitation in existing benchmarks that rely solely on structure-only function call graphs. To address this, the authors propose a semantic enrichment framework that augments graph topology with function-level attributes and LLM-based code embeddings, demonstrating that this data-centric approach improves robustness under distribution shift and complements model-based methods.
By Ngoc N. Tran, Anwar Said, Waseem Abbas, Tyler Derr, Xenofon D. Koutsoukos