arXiv AI

Dimensionality Reduction for Cyberattack Classification: A Comparative Evaluation of PCA and Linear Predictive Coding

arXiv:2606. 05584v1 Announce Type: cross Abstract: High-dimensional feature representations are widely used in machine learning-based cyberattack detection systems.

arXiv Machine Learning
Sep 24

Enhancing Multiclass Malware Classification in Resource-Constrained Environments

The paper presents a lightweight machine‑learning approach for multi‑class malware detection on resource‑constrained devices. Using a LightGBM classifier with SMOTE oversampling, SOM‑US undersampling, and Genetic‑Algorithm feature selection, the authors achieve 89.1 % accuracy on four malware families and 76 % on 16 individual malware types. A second Random‑Forest model further improves family classification to 91.2 % and individual classification to 78.7 %.

By Abdul Khalek Alve, Alif Rahman, Saadman Zaman, Sazzad Hossen Himel, Muhammad Iqbal Hossain
arXiv Machine Learning
Jun 5

Hybrid CNN-LSTM Framework for Intelligent Cyber Attack Detection and Prevention in U.S. Critical Digital Infrastructure: A Comparative Machine Learning Evaluation on CSE-CIC-IDS2018

arXiv:2606. 05714v1 Announce Type: cross Abstract: Digital infrastructure is growing at a rapid pace in the United States, and as a result, exposure to advanced cyber threats to critical sectors including healthcare, finance, transportation, energy and government systems is growing.

By Md. Iqbal Hossan, Md. Serajul Kabir Chowdhury Rubel, Md. Arifur Rahman, B. M. Taslimul Haque
arXiv Machine Learning
Sep 21

An Introduction to Compression-Based Machine Learning

The paper discusses how any lossless compression algorithm can be transformed into a machine learning method using Normalized Compression Distance or the Minimum Description Length principle, and conversely how any auto‑regressive model can become a lossless compressor via entropy coding. It surveys and formalizes these strategies, introduces a design framework for compression‑based ML, and empirically validates that such methods can match conventional baselines and outperform them on malware detection, achieving accuracy gains up to 0.62 by varying design choices.

By John Hurwitz, Edward Raff, Charles K. Nicholas