arXiv Machine Learning

On the Impact of Entropy-based Features

arXiv:2607. 15379v1 Announce Type: cross Abstract: Network anomaly detection is increasingly challenging due to the growing diversity and variability of traffic patterns, which are not always well captured by traditional statistical features.

Hugging Face Trending Papers
Jun 29

Multi-Level Distributional Entropy for Explainable Network Intrusion Detection

Machine learning network intrusion detection systems (IDS) rely on aggregate flow statistics that discard distributional structure, while established entropy measures require raw packet sequences unavailable in pre-aggregated flow datasets. We propose Multi-Level Distributional Entropy (MDE), an analytical framework that derives interpretable entropy features directly from flow-level summary statistics at three levels: within-flow Gaussian differential entropy, cross-directional Jensen-Shannon divergence (JSD), and Transmission Control Protocol (TCP) flag-pattern Shannon entropy, without raw packet access or training data.

arXiv AI
Jun 30

Multi-Level Distributional Entropy for Explainable Network Intrusion Detection

arXiv:2606. 29797v1 Announce Type: cross Abstract: Machine learning network intrusion detection systems (IDS) rely on aggregate flow statistics that discard distributional structure, while established entropy measures require raw packet sequences unavailable in pre-aggregated flow datasets.

By Mohamed Aly Bouke, Md Shohel Sayeed, Swee-Huay Heng, Azizol Abdullah, Mohamed Othman
arXiv Machine Learning
Sep 16

Beyond Measurement Metrics: A Human-Centered Framework for Semantic Validation of Network Traffic Classification

The paper proposes a human-centered framework for validating the semantic soundness of machine learning models used in network traffic classification. It extends existing knowledge-generation methods by integrating data, models, explainability tools, visualizations, and expert reasoning to iteratively explore, verify, and refine model behavior and preprocessing steps. The framework is built on literature findings, benchmark analyses, XAI experience, and expert feedback, offering practical guidance for ensuring models learn trustworthy, semantically meaningful patterns rather than spurious correlations.

By Igor Cherepanov, David Sessler, Alex Ulmer, Thorsten May, J\"orn Kohlhammer
arXiv AI
Sep 24

Comparative Evaluation of Static Embedding Models for HTTP Request Anomaly Detection

The paper benchmarks static embedding models—Word2Vec, FastText, and Doc2Vec—for detecting anomalous HTTP requests using a single‑class classification framework. It introduces HEDA, a modular pipeline that trains both embeddings and detectors solely on benign traffic in an unsupervised setting. Experiments on synthetic and real datasets show that FastText embeddings consistently yield high detection rates with controlled false positives.

By Amanda Riverol, Gustavo Betarte, Rodrigo Mart\'inez, \'Alvaro Pardo
arXiv AI
2d ago

Jev-IDS: System One Models for Network Intrusion Detection

JEV-IDS is an open experimental general network intrusion detection system that uses the Jev System One Model to detect zero‑day intrusions even when labeled data are scarce. The system processes one flow per request and asks the model two questions: a binary attack probability and a finite‑choice traffic category. In tests on a 300‑flow NSL‑KDD pilot split, JEV-IDS achieved an F1‑score of 0.859, precision of 0.941, recall of 0.790, and a novel‑attack recall of 0.838, while being 4.8 times faster and 3.8 times cheaper than GPT‑5.6 Luna and producing 15 times fewer false alarms than a low‑data Random Forest.

By Paulo Severo, Silvio E. Quincozes, Amanda Dias