arXiv Machine Learning

EXE-Bench: Ranking the Tradeoffs of AI-based Windows Malware Detectors for Real-World Usability

arXiv:2607. 24177v1 Announce Type: cross Abstract: Due to the lack of systematic evaluations, we are not yet able to determine which AI-based Windows malware detector to deploy in production, since existing evaluations (i) differ in terms of data used for both training and testing; (ii) do not consider temporal analysis to showcase whether models withstand the passage of time; (iii) avoid security evaluations with adversarial attacks that could highlight their brittleness against content-injection attacks; and (iv) neglect the computational requirements for deployment, risking slow inference on endpoints.

arXiv Machine Learning
Sep 18

Delphi Scanner: efficient and interpretable static malware detection via API sequence modeling

Delphi Scanner is a static malware detection system for Windows PE files that balances efficiency and interpretability. It employs a convolutional neural network to model Windows API sequences and a rule‑based interpretation layer to map APIs to high‑level malicious capabilities. Tested on over 190,000 PE files, it achieves 95.35% accuracy with a 1.53 MB model, and demonstrates robustness against out‑of‑distribution samples and adversarial manipulations.

By Bijied Brahimi, Vincent Cohadon, Gabriel Glazman, Rayan Al Mohaize, Omran Berjawi, Rida Khatoun
arXiv AI
Sep 10

SCRIPTIOC-BENCH: A Benchmark for Recognizing Actionable Threat Intelligence from Script-Based Malware using LLMs

SCRIPTIOC-BENCH is a benchmark designed to evaluate how well large language models can statically extract indicators of compromise (IOCs) from script-based malware. It contains 634 manually verified JavaScript, PowerShell, and VBScript samples and covers four IOC types—URLs, domains, IP addresses, and filesystem artifacts—while distinguishing between directly exposed and encoded indicators. Experiments show that even the best models achieve only 65.4 F1, and a false‑positive taxonomy is introduced to analyze error patterns, with two mitigations (deterministic string utilities and task‑specific adaptation) improving precision and shifting errors toward sample‑grounded mismatches.

By Hanna Kim, Jian Cui, Minkyoo Song, Hwanjo Heo, Seungwon Shin, Kimin Lee, Xiaojing Liao
arXiv AI
Jun 4

CyberGym-E2E: Scalable Real-World Benchmark for AI Agents' End-to-End Cybersecurity Capabilities

arXiv:2606. 04460v1 Announce Type: cross Abstract: AI has the potential to transform cybersecurity by enabling systems that can autonomously detect, analyze, and remediate software vulnerabilities.

By Tianneng Shi, Robin Rheem, Dongwei Jiang, Mona Wang, Francisco De La Riega, Zhun Wang, Jingzhi Jiang, Alexander Cheung, Sean Tai, Jonah Cha, Jianhong Tu, Gabriel Han, Chenguang Wang, Jingxuan He, Wenbo Guo, Dawn Song
arXiv Machine Learning
Aug 31

REPLICANT: Learning Policies for Evading and Hardening Malware Detectors

The paper introduces Replicant, a deep reinforcement learning framework that learns to evade malware detectors under a strict label‑only black‑box threat model. Replicant generates reusable policies for modifying malware samples and deciding when to query the target, and it transfers across different samples, detectors, and feature spaces. In experiments on seven Android malware detectors and three feature spaces, Replicant achieves a mean attack success rate of 78.8%, outperforming state‑of‑the‑art methods by 20.9%–39.2% and providing a stronger signal for adversarial training to harden detectors.

By Shae McFadden, Ilias Tsingenopoulos, Mario D'Onghia, Alexander Herzog, Myles Foley, Chris Hicks, Lorenzo Cavallaro, Fabio Pierazzi