The paper introduces a leave‑a‑window‑out estimator for next‑token functionals, such as the surprise probability and test error, in sequences of random variables. By deleting a window of length τ after each index, the estimator generalizes leave‑one‑out and achieves parametric error decay for stationary β‑mixing processes that admit a Marton coupling. The authors provide both upper bounds and a minimax lower bound for the surprise probability, and demonstrate through simulations that their method outperforms traditional baselines on Markov, moving‑average, and autoregressive processes.
By Milind Nakul, Vidya Muthukumar, Ashwin Pananjady
arXiv:2608. 25551v1 Announce Type: new Abstract: Stochastic gradient descent (SGD) is typically analyzed at a deterministic horizon chosen before the algorithm is run, even though practical stopping decisions are made adaptively by inspecting the evolving trajectory.
By Liviu Aolaritei, Lucas L\'evy, Francis Bach, Michael I. Jordan
arXiv:2608. 07224v1 Announce Type: new Abstract: Motivated by LLMs, which generate outputs by iteratively sampling from next-token distributions, we introduce a PAC-learning model for binary stochastic autoregressive learning.
By Ilan Doron-Arad, Idan Mehalel, Elchanan Mossel
arXiv:2608. 14408v1 Announce Type: cross Abstract: We study online statistical inference for functionals of the return distribution under a fixed policy.
By Yang Peng, Liangyu Zhang
The paper extends the idea that contexts are cheap for linear bandits from i.i.d. settings to Markovian context processes. By assuming uniform geometric ergodicity, the authors construct a stationary surrogate action set and use a delayed‑update scheme to mitigate bias from nonstationary conditional context distributions. They provide a phased algorithm for unknown stationary distributions and achieve high‑probability regret bounds comparable to standard linear bandit oracles in fast‑mixing regimes, with empirical validation showing gains over LinUCB.
By Kaan Buyukkalayci, Osama Hanna, Christina Fragouli
arXiv:2602. 05657v2 Announce Type: replace Abstract: The study of tail behaviour of SGD-induced processes has been attracting a lot of interest, due to offering strong guarantees with respect to individual runs of an algorithm.
By Aleksandar Armacki, Dragana Bajovi\'c, Du\v{s}an Jakoveti\'c, Soummya Kar, Ali H. Sayed
arXiv:2603. 06957v2 Announce Type: replace-cross Abstract: We study post-training linear autoregressive models with outcome and process rewards.
By Alireza Mousavi-Hosseini, Murat A. Erdogdu
arXiv:2609.38524v1 Announce Type: new
Abstract: We consider estimating the one-step-ahead conditional distribution of a multivariate stochastic process. Many existing approaches rely on assumptions s...
By Michael Wieck-Sosa, Cosma Rohilla Shalizi
arXiv:2502. 09884v4 Announce Type: replace-cross Abstract: We consider linear two-time-scale stochastic approximation algorithms driven by martingale noise.
By Seo Taek Kong, Sihan Zeng, Thinh T. Doan, R. Srikant
The paper establishes a shrinking‑tube concentration bound for projected stochastic approximation driven by an adaptive Markov chain, guaranteeing that after a chosen time every iterate stays within a tolerance that tightens over time. The bound’s probability of any exit after that time decays polynomially, and a matching lower bound shows this exponent is optimal under finite second moments. Extensions to recursions with martingale‑difference noise and predictable bias reveal how noise scale and bias affect exit‑probability decay and tube shrinkage, with applications to inventory learning and numerical gradient accuracy.
By Jin Li, Ye Luo, Xiaowei Zhang
arXiv:2606. 11149v1 Announce Type: new Abstract: We study the problem of learning a drifting concept in the presence of Massart noise.
By Mingchen Ma, Guyang Cao, Jelena Diakonikolas, Ilias Diakonikolas
arXiv:2605. 26000v2 Announce Type: replace-cross Abstract: Stochastic gradient descent (SGD) is foundational to large-scale statistical learning and stochastic optimization.
By Jose Blanchet, Peter Glynn, Wenhao Yang