arXiv Machine Learning

Stream Assembly Is an Uncontrolled Treatment in Streaming Intrusion-Detection Benchmarks

The paper demonstrates that the way evaluation streams are assembled in streaming intrusion‑detection benchmarks—by interleaving, pooling, or replaying network captures—acts as an uncontrolled experimental variable that can significantly alter performance metrics. In the CICIDS2017 benchmark, reordering the same set of records under a fixed split changes the held‑out samples’ overlap, prevalence, and even reverses the ranking of two deterministic scorers. Similar effects are observed in the LITNET‑2020 benchmark, where pooling disjoint captures yields a single operating point that masks large variations in per‑capture prevalences, and minor changes in batch composition can shift reported AUC‑PR values by a few thousandths.

arXiv Machine Learning
Jun 26

CALIBURN: Operationally Calibrated Streaming Intrusion Detection with Regime-Dependent Conformal Risk Control

arXiv:2605. 24696v2 Announce Type: replace-cross Abstract: Streaming intrusion detection systems must process flows continuously under bounded memory, yet most leave alerting-threshold selection as a post-hoc tuning problem incompatible with production, where operators commit in advance to alert budgets, misclassification costs, and Service Level Objectives.

By Michel A. Youssef
arXiv Machine Learning
Sep 7

Candidate Comparability Before Promotion: Conditional Validation in Adaptive Network Intrusion Detection

The paper investigates how to properly validate candidate models before promoting them to replace incumbent classifiers in adaptive network intrusion detection systems. It demonstrates that promotion decisions can be biased by how challengers are constructed and the amount of evidence they receive, and that using self‑contained challenger pipelines and sufficient candidate evidence reduces apparent promotion harm. The study also shows that policy rankings shift with candidate comparability and that no single update policy dominates across benchmarks.

By Roberto Fern\'andez-Barrios, Iker Pastor-L\'opez, Amaia Pikatza-Huerga, Pablo Garc\'ia Bringas