arXiv Machine Learning

When the Martingale Never Stops Firing: Anytime-Valid Gating on Real Forecast Streams

arXiv Machine Learning
Aug 4

Real-Time Detection and Repair of LLM Agent Failures

arXiv:2608. 02464v1 Announce Type: cross Abstract: LLM agents fail mid-episode -- they loop, cascade tool errors, drift off goal, fabricate results, or silently absorb corrupted content -- and the standard remedy, judging every step with a second LLM, costs more than the agent itself.

By Sunny Dubey
arXiv Machine Learning
Sep 22

Information-Geometric First-Passage Monitoring of Distributional Stability in Stochastic Systems

The paper presents a runtime monitoring framework for stochastic systems that distinguishes normal distributional relaxation from regime changes while limiting false alarms. It combines relative‑entropy dissipation, information geometry, and sequential inference within a bounded first‑passage architecture, employing Gaussian window surrogates, covariance shrinkage, and conformal ranking aggregated by a mixture power‑martingale. Validation on Ornstein–Uhlenbeck dynamics and network intrusion datasets (NSL‑KDD, UNSW‑NB15) shows high detection rates with low false positives, highlighting calibration transport as a key deployment challenge.

By Hikmat Karimov, Rahid Zahid Alekberli
arXiv Machine Learning
Sep 17

Capability Emergence Can Be Forecast: Per-Seed, In Advance, With Calibrated Intervals, Certified False Alarms, and a Blind Pre-Registered Gate

The paper demonstrates that emergent capabilities in machine learning models can be forecasted with lead time, calibrated uncertainty, and controlled false‑alarm rates. Using per‑seed analysis on transformers, the authors show that the formation time of a previous‑token head predicts the emergence of an induction head with Spearman ρ = 0.977 and a median lead of 975 training steps. Conformal intervals, blind pre‑registered tests, and a multiplicative rule relating anchor and event times further validate the predictive framework across multiple model families and configurations.

By Gunner Levi Howe
arXiv Machine Learning
Jul 13

Global Sequential Testing for Multi-Stream Auditing

arXiv:2602. 21479v3 Announce Type: replace-cross Abstract: Across many risk-sensitive areas, it is critical to continuously audit machine learning systems as we receive more data to quickly determine if they are performing as designed.

By Beepul Bharti, Ambar Pal, Jeremias Sulam