arXiv Machine Learning
Sep 18

Next-token functional estimation

The paper introduces a leave‑a‑window‑out estimator for next‑token functionals, such as the surprise probability and test error, in sequences of random variables. By deleting a window of length τ after each index, the estimator generalizes leave‑one‑out and achieves parametric error decay for stationary β‑mixing processes that admit a Marton coupling. The authors provide both upper bounds and a minimax lower bound for the surprise probability, and demonstrate through simulations that their method outperforms traditional baselines on Markov, moving‑average, and autoregressive processes.

By Milind Nakul, Vidya Muthukumar, Ashwin Pananjady
arXiv Machine Learning
Aug 10

Stochastic Autoregressive Learning

arXiv:2608. 07224v1 Announce Type: new Abstract: Motivated by LLMs, which generate outputs by iteratively sampling from next-token distributions, we introduce a PAC-learning model for binary stochastic autoregressive learning.

By Ilan Doron-Arad, Idan Mehalel, Elchanan Mossel
arXiv Machine Learning
Sep 10

Mixing Makes Markovian Contexts Cheap for Linear Bandits

The paper extends the idea that contexts are cheap for linear bandits from i.i.d. settings to Markovian context processes. By assuming uniform geometric ergodicity, the authors construct a stationary surrogate action set and use a delayed‑update scheme to mitigate bias from nonstationary conditional context distributions. They provide a phased algorithm for unknown stationary distributions and achieve high‑probability regret bounds comparable to standard linear bandit oracles in fast‑mixing regimes, with empirical validation showing gains over LinUCB.

By Kaan Buyukkalayci, Osama Hanna, Christina Fragouli