arXiv Machine Learning By Yang Peng, Liangyu Zhang

Online Inference in Distributional Temporal-Difference Learning

Read the original on arXiv Machine Learning →

arXiv:2608. 14408v1 Announce Type: cross Abstract: We study online statistical inference for functionals of the return distribution under a fixed policy.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 8

Model-based Bootstrap of Controlled Markov Chains

arXiv:2605. 12410v2 Announce Type: replace-cross Abstract: We propose and analyze a model-based bootstrap for transition kernels in finite controlled Markov chains (CMCs) with possibly nonstationary or history-dependent control policies, a setting that arises naturally in offline reinforcement learning (RL) when the behavior policy generating the data is unknown.

By Ziwei Su, Imon Banerjee, Diego Klabjan