arXiv:2608.30502v1 Announce Type: new
Abstract: Machine learning systems are increasingly corrected while they run, and the decision of when to intervene is increasingly delegated to statistical moni...
By Weijia Han, Lisha Qu
arXiv:2504. 19952v2 Announce Type: replace-cross Abstract: We present two general lower bounds for stopping times of sequential tests between arbitrary composite nulls $\mathcal P$ and alternatives $\mathcal Q$.
By Shubhada Agrawal, Ashwin Ram, Aaditya Ramdas
arXiv:2602. 13848v2 Announce Type: replace Abstract: We propose a sequential test for detecting arbitrary distribution shifts that allows conformal test martingales (CTMs) to work under a fixed, reference-conditional setting.
By Shalev Shaer, Yarin Bar, Drew Prinster, Yaniv Romano
arXiv:2606. 30338v1 Announce Type: new Abstract: External evaluations are becoming increasingly central to the governance of AI systems.
By Ioannis Pitsiorlas, Martha V. Sourla, Marios Kountouris
arXiv:2603. 17925v2 Announce Type: replace-cross Abstract: We consider a variant of sequential testing by betting where, at each time step, the statistician is presented with multiple data sources (arms) and obtains data by choosing one of the arms.
By Ricardo J. Sandoval, Ian Waudby-Smith, Michael I. Jordan
arXiv:2609.37841v1 Announce Type: new
Abstract: Masked generative models offer parallel token prediction, but accurate parallel sampling must account for dependencies among tokens. When dependencies...
By Ryotaro Kawata, Satoshi Hayakawa, Taiji Suzuki
arXiv:2607. 11653v1 Announce Type: new Abstract: Black-box conditional quantile forecasts are widely used for sequential decisions under asymmetric costs, such as inventory planning in supply chain management.
By Ivane Antonov, Sohom Mukherjee, Richard Pibernik, Yo Joong Choe
arXiv:2609.38695v1 Announce Type: cross
Abstract: Generative AI has dramatically accelerated the rate at which new treatments---from novel pharmaceuticals to online marketing campaigns---can be conce...
By Ricardo J. Sandoval, David Arbour, Avi Feller, Michael I. Jordan
arXiv:2606. 20859v2 Announce Type: replace-cross Abstract: A fundamental assumption in statistics and machine learning is that ``the future looks like the past,'' formalized as exchangeability: the joint data distribution is order-invariant.
By Johan Hallberg Szabadv\'ary
arXiv:2606. 15237v1 Announce Type: cross Abstract: Ensemble classifiers are predictive models that combine the results of simpler base models, often by majority vote.
By Joseph Kalman, Amit Moscovich
The paper proves that in time‑series validation three desirable properties—training sufficiency, test coverage, and temporal causality—cannot all be satisfied simultaneously. It introduces quantitative bounds involving the smallest training fraction (α), test coverage (β), future training fraction (Λ), and distance to nearest future training point (δ), showing that exceeding the causal frontier α+β=1 requires training on future data that must lie within (1−α)T of a test point. The authors demonstrate that the impact of such future leakage depends on distance rather than amount, and compare different validation schemes (walk‑forward, k‑fold, purged k‑fold) in terms of their position on this Pareto frontier, illustrating the trade‑offs with empirical results on noise data.
By Jiayu Li
The paper introduces a leave‑a‑window‑out estimator for next‑token functionals, such as the surprise probability and test error, in sequences of random variables. By deleting a window of length τ after each index, the estimator generalizes leave‑one‑out and achieves parametric error decay for stationary β‑mixing processes that admit a Marton coupling. The authors provide both upper bounds and a minimax lower bound for the surprise probability, and demonstrate through simulations that their method outperforms traditional baselines on Markov, moving‑average, and autoregressive processes.
By Milind Nakul, Vidya Muthukumar, Ashwin Pananjady