arXiv Machine Learning By Michel A. Youssef

CALIBURN: Operationally Calibrated Streaming Intrusion Detection with Regime-Dependent Conformal Risk Control

Read the original on arXiv Machine Learning →

arXiv:2605. 24696v2 Announce Type: replace-cross Abstract: Streaming intrusion detection systems must process flows continuously under bounded memory, yet most leave alerting-threshold selection as a post-hoc tuning problem incompatible with production, where operators commit in advance to alert budgets, misclassification costs, and Service Level Objectives.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 7

Candidate Comparability Before Promotion: Conditional Validation in Adaptive Network Intrusion Detection

The paper investigates how to properly validate candidate models before promoting them to replace incumbent classifiers in adaptive network intrusion detection systems. It demonstrates that promotion decisions can be biased by how challengers are constructed and the amount of evidence they receive, and that using self‑contained challenger pipelines and sufficient candidate evidence reduces apparent promotion harm. The study also shows that policy rankings shift with candidate comparability and that no single update policy dominates across benchmarks.

By Roberto Fern\'andez-Barrios, Iker Pastor-L\'opez, Amaia Pikatza-Huerga, Pablo Garc\'ia Bringas
arXiv AI
Sep 18

Refuse, Decompose, Refresh: A Claim-Safe Protocol for Closed-Loop AI Evaluation

The paper introduces a claim‑safe protocol for evaluating closed‑loop AI systems, consisting of three actions: Refuse, Decompose, and Refresh. It demonstrates the protocol in a simulator with 24 policy components and 1,440 held‑out cases, showing that abstention and stable false admission rates are low while providing detailed statistical diagnostics. The approach emphasizes that evaluation results should be tied to observable support and statistical calibration rather than a single PASS/FAIL label.

By Peiying Zhu, Sidi Chang