arXiv AI

Improving IoT Intrusion Detection Through SMOTE-Based Oversampling and Extended Multi-Model Evaluation on Side-Channel Power Data

arXiv:2606. 00161v1 Announce Type: cross Abstract: The detection of intrusions in IoT-based networks poses challenges that cannot be overcome using traditional machine learning methods.

arXiv Machine Learning
Sep 25

Unmasking Shortcut Learning in IoT Intrusion Detection: A Forensic, Multi-Paradigm Evaluation of Feature Dependence and Data Leakage

The paper investigates whether machine learning models for IoT intrusion detection truly learn attack patterns or rely on dataset shortcuts. Using the CyberFlowIoT-GICAP benchmark, the authors evaluate four learning paradigms across different feature sets and split strategies, finding that performance is largely driven by feature representation and that tree-based models can exploit temporal artifacts. The study also highlights asymmetric attack detectability and proposes a four-point protocol checklist for realistic evaluation.

By Uday Shankar Roy, Mahbuba Jahan Minu
arXiv AI
Sep 7

Data-Optimized Contingency Screening: A Machine Learning Approach to Power System Security

The paper presents a machine‑learning framework for classifying power‑system contingencies into safe, moderate, or severe categories. Using Newton‑Raphson load flow data, the study applies SMOTE, PCA, and classifiers (KNN, Random Forest, SVM) to IEEE‑14 and IEEE‑30 bus systems, evaluating performance with precision, recall, and F1 score. Random Forest achieved the highest F1 scores, while PCA improved overall performance more than SMOTE, which boosted recall at the cost of some false positives.

By Joshua Salako, Folajimi Osikomaiya, Olakorede Olamiju
arXiv Machine Learning
Sep 24

Enhancing Multiclass Malware Classification in Resource-Constrained Environments

The paper presents a lightweight machine‑learning approach for multi‑class malware detection on resource‑constrained devices. Using a LightGBM classifier with SMOTE oversampling, SOM‑US undersampling, and Genetic‑Algorithm feature selection, the authors achieve 89.1 % accuracy on four malware families and 76 % on 16 individual malware types. A second Random‑Forest model further improves family classification to 91.2 % and individual classification to 78.7 %.

By Abdul Khalek Alve, Alif Rahman, Saadman Zaman, Sazzad Hossen Himel, Muhammad Iqbal Hossain
arXiv AI
Jun 9

SHIELD-IDS: Structurally Heterogeneous Ensemble with Integrated Layered Defense for Intrusion Detection Systems

arXiv:2606. 07716v1 Announce Type: cross Abstract: Adversarial attacks pose a serious and growing threat to Machine Learning (ML)-based Intrusion Detection Systems (IDS), where imperceptible perturbations to network flow features can systematically mislead classifiers into accepting malicious traffic as benign.

By Maryam Zaman, Muhammad Khuram Shahzad
arXiv AI
Jul 21

Data Balancing Strategies: A Systematic Survey of Resampling and Augmentation Methods

arXiv:2505. 13518v3 Announce Type: replace-cross Abstract: Imbalanced datasets, where one class significantly outnumbers others, remain a persistent challenge in machine learning, often biasing predictions toward the majority class and degrading classifier performance.

By Behnam Yousefimehr, Mehdi Ghatee, Javad Fazli, Shervin Ghaffari, Zahra Rafei, Mohammad Amin Seifi, Sajed Tavakoli, Abolfazl Nikahd, Mahdi Razi Gandomani, Alireza Orouji, Ramtin Mahmoudi Kashani, Sarina Heshmati, Negin Sadat Mousavi