arXiv Machine Learning

Candidate Comparability Before Promotion: Conditional Validation in Adaptive Network Intrusion Detection

The paper investigates how to properly validate candidate models before promoting them to replace incumbent classifiers in adaptive network intrusion detection systems. It demonstrates that promotion decisions can be biased by how challengers are constructed and the amount of evidence they receive, and that using self‑contained challenger pipelines and sufficient candidate evidence reduces apparent promotion harm. The study also shows that policy rankings shift with candidate comparability and that no single update policy dominates across benchmarks.

arXiv Machine Learning
Jun 26

CALIBURN: Operationally Calibrated Streaming Intrusion Detection with Regime-Dependent Conformal Risk Control

arXiv:2605. 24696v2 Announce Type: replace-cross Abstract: Streaming intrusion detection systems must process flows continuously under bounded memory, yet most leave alerting-threshold selection as a post-hoc tuning problem incompatible with production, where operators commit in advance to alert budgets, misclassification costs, and Service Level Objectives.

By Michel A. Youssef
arXiv Machine Learning
Jul 2

Forensic-Oriented Intrusion Detection Using Synthetic Network Traffic Data and Explainable Artificial Intelligence

arXiv:2607. 00763v1 Announce Type: cross Abstract: Digital forensic investigations of network intrusions require analytical outputs that are traceable, reproducible, and court-defensible - requirements existing machine learning pipelines do not satisfy, since they treat original evidence as training data and produce opaque classifications without instance-level justification.

By Jose Luis Vela Alonso, Carmen Pellicer
arXiv AI
1d ago

HydroJEV: A one-second, training-free screen for cyber-attack and fault attribution in water distribution networks

HydroJEV is a training‑free, one‑second model that classifies SCADA alarms in water distribution networks into cyberattack, physical fault, normal transient, or faulty sensor. In a benchmark on the C‑Town EPANET network, HydroJEV matched a hand‑written rule tree and outperformed a supervised classifier, especially when few labeled events were available, while being 20‑40 times faster than large language models. When combined with a rule tree gate, it reduced the need for human review by about a third without sacrificing accuracy.

By Tianwei Mu, Shengyan Jiang, Mingzhe Yuan, Qing Luo, Min Xiao, Wenhong Wang, Jun Li, Manhong Huang
arXiv AI
Sep 4

Privacy, Robustness, and Fairness Trade-offs in Federated Intrusion Detection: Geometric Indistinguishability at the Aggregation Interface

The paper investigates how privacy guarantees, robustness to Byzantine attacks, and detection coverage for rare intrusion types interact in federated network intrusion detection systems. It introduces geometric indistinguishability to explain how privacy noise can obscure minority-class signals, and demonstrates through experiments on UNSW‑NB15 that combining differential privacy with robust aggregation can disproportionately harm detection of rare attacks. The study highlights that these properties cannot be treated as independently composable and calls for aggregation‑aware modeling and sample‑aware evaluation to build trustworthy federated NIDS.

By Adrita Rahman Tory, ABM Shawkat Ali, Md Abu Layek, Khondokar Fida Hasan
arXiv AI
Sep 15

A Three-Axis Stress Test of LLM vs Classical ML for Network Intrusion Detection under Distribution Shift and Adversarial Evasion

The study compares XGBoost and RoBERTa‑LoRA for network intrusion detection across three evaluation axes: same‑dataset performance, cross‑dataset transfer, and adversarial evasion. Both models perform similarly on the same dataset, but XGBoost outperforms RoBERTa‑LoRA by 15 F1 points and 25 balanced accuracy points when transferred to a different network, while RoBERTa‑LoRA wins by about 17 F1 points under adversarial evasion. Feature‑leakage ablation shows that cross‑dataset transfer improvements are non‑monotonic and directional, suggesting leakage is spread across features rather than isolated. "whyItMatters":"The findings demonstrate that a model’s superiority depends on the specific robustness axis evaluated, underscoring the need for multi‑axis, multi‑metric testing in network intrusion detection research."

By Muhammad Ebad Atif, Muhammad Haider Ali