arXiv Machine Learning

Global Sequential Testing for Multi-Stream Auditing

arXiv:2602. 21479v3 Announce Type: replace-cross Abstract: Across many risk-sensitive areas, it is critical to continuously audit machine learning systems as we receive more data to quickly determine if they are performing as designed.

arXiv Machine Learning
Jun 5

Multi-Armed Sequential Hypothesis Testing by Betting

arXiv:2603. 17925v2 Announce Type: replace-cross Abstract: We consider a variant of sequential testing by betting where, at each time step, the statistician is presented with multiple data sources (arms) and obtains data by choosing one of the arms.

By Ricardo J. Sandoval, Ian Waudby-Smith, Michael I. Jordan
arXiv AI
3d ago

Always-On Experimentation

arXiv:2609.38695v1 Announce Type: cross Abstract: Generative AI has dramatically accelerated the rate at which new treatments---from novel pharmaceuticals to online marketing campaigns---can be conce...

By Ricardo J. Sandoval, David Arbour, Avi Feller, Michael I. Jordan
arXiv Machine Learning
Sep 25

The Impossible Trinity of Time-Series Validation: A Conservation Law among Training Sufficiency, Test Coverage, and Temporal Causality

The paper proves that in time‑series validation three desirable properties—training sufficiency, test coverage, and temporal causality—cannot all be satisfied simultaneously. It introduces quantitative bounds involving the smallest training fraction (α), test coverage (β), future training fraction (Λ), and distance to nearest future training point (δ), showing that exceeding the causal frontier α+β=1 requires training on future data that must lie within (1−α)T of a test point. The authors demonstrate that the impact of such future leakage depends on distance rather than amount, and compare different validation schemes (walk‑forward, k‑fold, purged k‑fold) in terms of their position on this Pareto frontier, illustrating the trade‑offs with empirical results on noise data.

By Jiayu Li
arXiv Machine Learning
Sep 18

Next-token functional estimation

The paper introduces a leave‑a‑window‑out estimator for next‑token functionals, such as the surprise probability and test error, in sequences of random variables. By deleting a window of length τ after each index, the estimator generalizes leave‑one‑out and achieves parametric error decay for stationary β‑mixing processes that admit a Marton coupling. The authors provide both upper bounds and a minimax lower bound for the surprise probability, and demonstrate through simulations that their method outperforms traditional baselines on Markov, moving‑average, and autoregressive processes.

By Milind Nakul, Vidya Muthukumar, Ashwin Pananjady