arXiv Machine Learning

A Lightweight Hybrid MLP-Based Framework for Real-Time Phishing URL Detection Using Structural URL Features

arXiv:2606. 00889v1 Announce Type: cross Abstract: Phishing attacks remain a major cybersecurity threat, exploiting deceptive URLs to steal sensitive user information.

arXiv AI
Sep 10

CoGReV: A Confidence-Gated Post-Hoc Non-Monotonic Belief Revision Framework for Phishing Website Classification

CoGReV is a hybrid framework that enhances machine‑learning phishing classifiers with a post‑hoc, non‑monotonic reasoning layer written in Answer Set Programming. It uses a confidence‑gated defeasible rule to revise low‑confidence phishing predictions toward legitimate only when website metadata is available, thereby allocating uncertain decisions to the reasoning layer while leaving confident ones to the classifier. The gated rule reduces false positives by 0.27 % of decisions and maintains recall within 0.7 % of the baseline, operating in linear time.

By Mainak Sen, Kumar Sankar Ray, Amlan Chakrabarti
arXiv Machine Learning
Sep 18

AURA: Adaptive Uncertainty-Routed Analysis for Email Threat Detection

AURA: Adaptive Uncertainty-Routed Analysis for Email Threat Detection is a multimodal system that evaluates both email content and embedded URLs to detect spam and phishing. It uses a two-layer approach: first, a URL classifier estimates prediction uncertainty, and only messages with high uncertainty are passed to a fine-tuned transformer encoder for deeper semantic analysis. Evaluated on eight diverse training corpora and two real-world datasets covering a decade of attacks, AURA achieves a macro F1-score of 0.9858 in-distribution and maintains scores above 0.94 on the NazPhish-Eval and GuenterTrap-Eval datasets, demonstrating strong generalization to new attack scenarios.

By Omran Berjawi, Walid fahs, Rida Khatoun
arXiv AI
Sep 24

Comparative Evaluation of Static Embedding Models for HTTP Request Anomaly Detection

The paper benchmarks static embedding models—Word2Vec, FastText, and Doc2Vec—for detecting anomalous HTTP requests using a single‑class classification framework. It introduces HEDA, a modular pipeline that trains both embeddings and detectors solely on benign traffic in an unsupervised setting. Experiments on synthetic and real datasets show that FastText embeddings consistently yield high detection rates with controlled false positives.

By Amanda Riverol, Gustavo Betarte, Rodrigo Mart\'inez, \'Alvaro Pardo
arXiv Machine Learning
Jun 5

Hybrid CNN-LSTM Framework for Intelligent Cyber Attack Detection and Prevention in U.S. Critical Digital Infrastructure: A Comparative Machine Learning Evaluation on CSE-CIC-IDS2018

arXiv:2606. 05714v1 Announce Type: cross Abstract: Digital infrastructure is growing at a rapid pace in the United States, and as a result, exposure to advanced cyber threats to critical sectors including healthcare, finance, transportation, energy and government systems is growing.

By Md. Iqbal Hossan, Md. Serajul Kabir Chowdhury Rubel, Md. Arifur Rahman, B. M. Taslimul Haque
arXiv Machine Learning
Sep 24

Enhancing Multiclass Malware Classification in Resource-Constrained Environments

The paper presents a lightweight machine‑learning approach for multi‑class malware detection on resource‑constrained devices. Using a LightGBM classifier with SMOTE oversampling, SOM‑US undersampling, and Genetic‑Algorithm feature selection, the authors achieve 89.1 % accuracy on four malware families and 76 % on 16 individual malware types. A second Random‑Forest model further improves family classification to 91.2 % and individual classification to 78.7 %.

By Abdul Khalek Alve, Alif Rahman, Saadman Zaman, Sazzad Hossen Himel, Muhammad Iqbal Hossain