arXiv Machine Learning

Information-Geometric First-Passage Monitoring of Distributional Stability in Stochastic Systems

The paper presents a runtime monitoring framework for stochastic systems that distinguishes normal distributional relaxation from regime changes while limiting false alarms. It combines relative‑entropy dissipation, information geometry, and sequential inference within a bounded first‑passage architecture, employing Gaussian window surrogates, covariance shrinkage, and conformal ranking aggregated by a mixture power‑martingale. Validation on Ornstein–Uhlenbeck dynamics and network intrusion datasets (NSL‑KDD, UNSW‑NB15) shows high detection rates with low false positives, highlighting calibration transport as a key deployment challenge.

arXiv AI
Jul 14

Calibrated e-CUSUM Decoding for Quantized Reasoning Models: Why Token Log-Probability Is the Wrong Observable for Decoding Monitors

arXiv:2607. 11317v1 Announce Type: new Abstract: Low-bit quantization makes small reasoning models inexpensive to deploy but can degrade their chains of thought.

By El Hassane Ettifouri (Novelis Research, Paris, France), Ayoub Belfatmi (Novelis Research, Paris, France), Mahaman Sanoussi Yahaya Alassan (Novelis Research, Paris, France), Walid Dahhane (Novelis Research, Paris, France)
Hugging Face Trending Papers
Jul 5

Asymptotic-Preserving A Posteriori Analysis of Diffusion and Flow-Matching Samplers

Diffusion and flow-matching samplers integrate a learned probability-flow ODE from a large noise scale down to a small terminal floor $σ_{\min}$, at which the score is stiff and the flow develops a boundary layer. We treat $σ_{\min}$ as a singular-perturbation parameter and determine which fixed-step samplers are asymptotic-preserving (AP), that is, stable and uniformly accurate as $σ_{\min}\to0$, casting the criteria as an a posteriori audit: residual functionals with $σ_{\min}$-uniform coefficients, computable on a pretrained checkpoint without ground-truth scores or exact trajectories.

arXiv AI
Sep 18

Refuse, Decompose, Refresh: A Claim-Safe Protocol for Closed-Loop AI Evaluation

The paper introduces a claim‑safe protocol for evaluating closed‑loop AI systems, consisting of three actions: Refuse, Decompose, and Refresh. It demonstrates the protocol in a simulator with 24 policy components and 1,440 held‑out cases, showing that abstention and stable false admission rates are low while providing detailed statistical diagnostics. The approach emphasizes that evaluation results should be tied to observable support and statistical calibration rather than a single PASS/FAIL label.

By Peiying Zhu, Sidi Chang
arXiv Machine Learning
Jun 26

CALIBURN: Operationally Calibrated Streaming Intrusion Detection with Regime-Dependent Conformal Risk Control

arXiv:2605. 24696v2 Announce Type: replace-cross Abstract: Streaming intrusion detection systems must process flows continuously under bounded memory, yet most leave alerting-threshold selection as a post-hoc tuning problem incompatible with production, where operators commit in advance to alert budgets, misclassification costs, and Service Level Objectives.

By Michel A. Youssef
arXiv Machine Learning
Sep 22

The Price of Self-Calibration: Exact Evidence Budgets and Manufactured Blind Sets in Adaptive Monitoring

The paper derives precise cost formulas for self‑calibrating monitors that adjust thresholds online to maintain a specified long‑run false‑alarm rate under arbitrary drift. It shows that the guarantee is an accounting identity, independent of the monitored signal, and provides exact evidence identities for both step and ramp drift scenarios, as well as an exact law for the fluctuation of the certificate’s own alarm rate. Additionally, it proves that any monitor designed to tolerate a drift class is blind to all faults in the difference of that class, identifying the blind set for speed‑bounded drift classes and quantifying power outside this set with a sharp Gaussian projection bound.

By Abdou-Raouf Atarmla