When the Martingale Never Stops Firing: Anytime-Valid Gating on Real Forecast Streams
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2505. 04608v5 Announce Type: replace-cross Abstract: Responsibly deploying artificial intelligence (AI) / machine learning (ML) systems in high-stakes settings arguably requires not only proof of system reliability, but also continual, post-deployment monitoring to quickly detect and address any unsafe behavior.
arXiv:2606. 19386v1 Announce Type: cross Abstract: Runtime monitors for autonomous agents commonly threshold an accumulated internal state - a behavioural baseline, a drift statistic, or, in our prior work, a modelled affective state.
arXiv:2608. 02464v1 Announce Type: cross Abstract: LLM agents fail mid-episode -- they loop, cascade tool errors, drift off goal, fabricate results, or silently absorb corrupted content -- and the standard remedy, judging every step with a second LLM, costs more than the agent itself.
arXiv:2607. 13048v1 Announce Type: cross Abstract: Streaming inference pipelines increasingly pair lightweight fast models with Large Language Models (LLMs) that provide rich semantic understanding at substantial cost.
The paper presents a runtime monitoring framework for stochastic systems that distinguishes normal distributional relaxation from regime changes while limiting false alarms. It combines relative‑entropy dissipation, information geometry, and sequential inference within a bounded first‑passage architecture, employing Gaussian window surrogates, covariance shrinkage, and conformal ranking aggregated by a mixture power‑martingale. Validation on Ornstein–Uhlenbeck dynamics and network intrusion datasets (NSL‑KDD, UNSW‑NB15) shows high detection rates with low false positives, highlighting calibration transport as a key deployment challenge.
The paper demonstrates that emergent capabilities in machine learning models can be forecasted with lead time, calibrated uncertainty, and controlled false‑alarm rates. Using per‑seed analysis on transformers, the authors show that the formation time of a previous‑token head predicts the emergence of an induction head with Spearman ρ = 0.977 and a median lead of 975 training steps. Conformal intervals, blind pre‑registered tests, and a multiplicative rule relating anchor and event times further validate the predictive framework across multiple model families and configurations.