arXiv Machine Learning By Andrea Ponte, Daniel Gibert, Matous Kozak, Dmitrijs Trizna, Maura Pintor, Battista Biggio, Fabio Roli, Luca Demetrio

EXE-Bench: Ranking the Tradeoffs of AI-based Windows Malware Detectors for Real-World Usability

Read the original on arXiv Machine Learning →

arXiv:2607. 24177v1 Announce Type: cross Abstract: Due to the lack of systematic evaluations, we are not yet able to determine which AI-based Windows malware detector to deploy in production, since existing evaluations (i) differ in terms of data used for both training and testing; (ii) do not consider temporal analysis to showcase whether models withstand the passage of time; (iii) avoid security evaluations with adversarial attacks that could highlight their brittleness against content-injection attacks; and (iv) neglect the computational requirements for deployment, risking slow inference on endpoints.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 18

Delphi Scanner: efficient and interpretable static malware detection via API sequence modeling

Delphi Scanner is a static malware detection system for Windows PE files that balances efficiency and interpretability. It employs a convolutional neural network to model Windows API sequences and a rule‑based interpretation layer to map APIs to high‑level malicious capabilities. Tested on over 190,000 PE files, it achieves 95.35% accuracy with a 1.53 MB model, and demonstrates robustness against out‑of‑distribution samples and adversarial manipulations.

By Bijied Brahimi, Vincent Cohadon, Gabriel Glazman, Rayan Al Mohaize, Omran Berjawi, Rida Khatoun
arXiv AI
Sep 10

SCRIPTIOC-BENCH: A Benchmark for Recognizing Actionable Threat Intelligence from Script-Based Malware using LLMs

SCRIPTIOC-BENCH is a benchmark designed to evaluate how well large language models can statically extract indicators of compromise (IOCs) from script-based malware. It contains 634 manually verified JavaScript, PowerShell, and VBScript samples and covers four IOC types—URLs, domains, IP addresses, and filesystem artifacts—while distinguishing between directly exposed and encoded indicators. Experiments show that even the best models achieve only 65.4 F1, and a false‑positive taxonomy is introduced to analyze error patterns, with two mitigations (deterministic string utilities and task‑specific adaptation) improving precision and shifting errors toward sample‑grounded mismatches.

By Hanna Kim, Jian Cui, Minkyoo Song, Hwanjo Heo, Seungwon Shin, Kimin Lee, Xiaojing Liao