arXiv Machine Learning

Online Inference in Distributional Temporal-Difference Learning

arXiv:2608. 14408v1 Announce Type: cross Abstract: We study online statistical inference for functionals of the return distribution under a fixed policy.

arXiv Machine Learning
Jul 8

Model-based Bootstrap of Controlled Markov Chains

arXiv:2605. 12410v2 Announce Type: replace-cross Abstract: We propose and analyze a model-based bootstrap for transition kernels in finite controlled Markov chains (CMCs) with possibly nonstationary or history-dependent control policies, a setting that arises naturally in offline reinforcement learning (RL) when the behavior policy generating the data is unknown.

By Ziwei Su, Imon Banerjee, Diego Klabjan
arXiv Machine Learning
Aug 19

A Finite-Iteration Theory for Asynchronous Categorical Distributional Temporal-Difference Learning

The paper investigates the finite‑iteration behavior of exact asynchronous recursions used in categorical distributional temporal‑difference (TD) learning. It analyzes both scalar categorical TD in the Cramér geometry and multivariate signed‑categorical TD in the maximum mean discrepancy geometry, showing that these methods can be viewed as single‑state stochastic‑approximation recursions that contract in a block‑supremum norm. The authors develop a restricted‑domain theory, derive discounted bounds under i.i.d. and Markovian sampling, and extend the analysis to undiscounted fixed‑horizon policy evaluation with horizon‑stacked categorical methods under episodic sampling, thereby providing a unified non‑asymptotic analysis across various settings.

By Ege C. Kaya, Abolfazl Hashemi
Hugging Face Trending Papers
Sep 17

Next-token functional estimation

Suppose we observe the first $n$ points of a sequence of random variables having length $n+1$, and wish to estimate a functional of the unobserved final point and the empirical measure of the $n$ obse...

arXiv Machine Learning
Sep 18

Next-token functional estimation

The paper introduces a leave‑a‑window‑out estimator for next‑token functionals, such as the surprise probability and test error, in sequences of random variables. By deleting a window of length τ after each index, the estimator generalizes leave‑one‑out and achieves parametric error decay for stationary β‑mixing processes that admit a Marton coupling. The authors provide both upper bounds and a minimax lower bound for the surprise probability, and demonstrate through simulations that their method outperforms traditional baselines on Markov, moving‑average, and autoregressive processes.

By Milind Nakul, Vidya Muthukumar, Ashwin Pananjady