arXiv Machine Learning

Classifier-Dependent Benefits of Pseudo-Labeling for Semi-Supervised Android Malware Attribution

The study evaluates pseudo‑labeling for semi‑supervised learning on Android malware attribution using six classifiers. Results show that the benefit of SSL varies strongly by classifier: SVM gains the most, LightGBM improves modestly, and Random Forest can be harmed at low label ratios. The approach particularly helps hard‑to‑classify families and achieves near‑optimal performance with about 800 labeled samples.

arXiv Machine Learning
Aug 5

ShielDroid: A Hybrid Approach Integrating Machine and Deep Learning for Android Malware Detection

arXiv:2608. 03250v1 Announce Type: cross Abstract: The rapid advancement of modern technology has led to a significant increase in the use of smart devices, such as smartphones and tablets, resulting in the widespread adoption of mobile applications.

By Md Faisal Ahmed, Zarin Tasnim Biash, Abu Raihan Shakil, Ahmed Ann Noor Ryen, Arman Hossain, Faisal Bin Ashraf, Muhammad Iqbal Hossain
arXiv Machine Learning
Sep 24

Enhancing Multiclass Malware Classification in Resource-Constrained Environments

The paper presents a lightweight machine‑learning approach for multi‑class malware detection on resource‑constrained devices. Using a LightGBM classifier with SMOTE oversampling, SOM‑US undersampling, and Genetic‑Algorithm feature selection, the authors achieve 89.1 % accuracy on four malware families and 76 % on 16 individual malware types. A second Random‑Forest model further improves family classification to 91.2 % and individual classification to 78.7 %.

By Abdul Khalek Alve, Alif Rahman, Saadman Zaman, Sazzad Hossen Himel, Muhammad Iqbal Hossain
arXiv AI
Aug 24

C-Score: Beyond Accuracy for Robustness Assessment in Semi-Supervised Learning under Open-World Unlabeled Contamination

The paper introduces C-Score, a diagnostic framework for evaluating pseudo‑label‑based semi‑supervised learning (SSL) when unlabeled data may contain out‑of‑distribution (OOD) samples. C-Score assesses training behavior across prediction, feature representation, and optimization, using metrics such as PLE, CCI, Sem‑Drift, and Grad‑Align. Experiments on CIFAR‑10 and CIFAR‑100 with various OOD sources show that C‑Score detects hidden degradation that clean accuracy alone fails to reveal, highlighting the need for internal diagnostic signals in SSL robustness assessment.

By Tsao-Lun Chen, Chi-Cheng Fu, Han-Yi E. Chou, Shun-Feng Su