Next-token functional estimation
Suppose we observe the first $n$ points of a sequence of random variables having length $n+1$, and wish to estimate a functional of the unobserved final point and the empirical measure of the $n$ obse...
The paper introduces a leave‑a‑window‑out estimator for next‑token functionals, such as the surprise probability and test error, in sequences of random variables. By deleting a window of length τ after each index, the estimator generalizes leave‑one‑out and achieves parametric error decay for stationary β‑mixing processes that admit a Marton coupling. The authors provide both upper bounds and a minimax lower bound for the surprise probability, and demonstrate through simulations that their method outperforms traditional baselines on Markov, moving‑average, and autoregressive processes.
Suppose we observe the first $n$ points of a sequence of random variables having length $n+1$, and wish to estimate a functional of the unobserved final point and the empirical measure of the $n$ obse...
arXiv:2608. 07224v1 Announce Type: new Abstract: Motivated by LLMs, which generate outputs by iteratively sampling from next-token distributions, we introduce a PAC-learning model for binary stochastic autoregressive learning.
arXiv:2608. 25551v1 Announce Type: new Abstract: Stochastic gradient descent (SGD) is typically analyzed at a deterministic horizon chosen before the algorithm is run, even though practical stopping decisions are made adaptively by inspecting the evolving trajectory.
arXiv:2608. 14408v1 Announce Type: cross Abstract: We study online statistical inference for functionals of the return distribution under a fixed policy.
arXiv:2602. 05657v2 Announce Type: replace Abstract: The study of tail behaviour of SGD-induced processes has been attracting a lot of interest, due to offering strong guarantees with respect to individual runs of an algorithm.
arXiv:2607. 19689v1 Announce Type: cross Abstract: We study the problem of recalibrating an online predictor [KE17, OKS24]: given an arbitrary "hint" sequence of forecasts, the learner must output new predictions that are calibrated while incurring small excess error relative to the original forecasts, under a proper loss.
arXiv:2605. 15806v2 Announce Type: replace Abstract: Neural operators excel as deterministic surrogates, but inevitably collapse to the conditional mean when applied to stochastic PDEs, discarding the variance and tail structure upon which uncertainty quantification depends.
arXiv:2603. 06957v2 Announce Type: replace-cross Abstract: We study post-training linear autoregressive models with outcome and process rewards.
The paper extends the idea that contexts are cheap for linear bandits from i.i.d. settings to Markovian context processes. By assuming uniform geometric ergodicity, the authors construct a stationary surrogate action set and use a delayed‑update scheme to mitigate bias from nonstationary conditional context distributions. They provide a phased algorithm for unknown stationary distributions and achieve high‑probability regret bounds comparable to standard linear bandit oracles in fast‑mixing regimes, with empirical validation showing gains over LinUCB.
arXiv:2601.05280v4 Announce Type: replace-cross Abstract: On the one hand, the question of whether Large Language Models (LLMs) are Solomonoff induction estimators has become an explicit question at...
arXiv:2607. 20309v1 Announce Type: cross Abstract: Covariate shift often occurs because, in many real applications, the source and the target observations may be generated from different distributions.
arXiv:2608. 14401v1 Announce Type: cross Abstract: In offline RL, estimating the optimal action-value function $Q^*$ can be formulated as solving the optimal Bellman equation based solely on offline observations.