arXiv:2502.07584v3 Announce Type: replace-cross
Abstract: Many learning algorithms can be represented as Markov processes, and understanding their generalization error is a central topic in learning...
By Benjamin Dupuis, Maxime Haddouche, George Deligiannidis, Umut Simsekli
The paper presents time‑uniform self‑normalized concentration bounds for stochastic processes in Hilbert spaces with vector‑valued noise, enabling regression‑error guarantees for both linear and nonlinear parametric operators. These results apply to possibly infinite‑dimensional inputs and outputs without requiring independence or mixing assumptions, and are derived in the context of sequentially collected, dependent data such as adaptive experimental design and dynamical‑system modelling.
By Rafael Oliveira
arXiv:2607. 23502v1 Announce Type: cross Abstract: We study empirical risk minimization for learning non-linear dynamical systems whose transition dynamics may switch over time.
By Sunny G. W. Wang, Hemant Tyagi
arXiv:2610. 01181v1 Announce Type: new Abstract: We consider stochastic games with independent controlled chains and unknown transition kernels, where players observe only their local states and realized payoffs.
By S. Rasoul Etesami
arXiv:2606. 16729v1 Announce Type: new Abstract: While there is an extensive body of work characterizing the sample complexity of discounted cumulative-reward MDPs, finite sample analyses for average-reward MDPs have been limited, and most existing works rely on restrictive assumptions such as ergodicity or access to a generative model.
By Jongmin Lee, Ernest K. Ryu, Vaneet Aggarwal
arXiv:2604.01024v2 Announce Type: replace
Abstract: We study model-based learning of finite-window policies in tabular partially observable Markov decision processes (POMDPs). A common approach to le...
By Philip Jordan, Maryam Kamgarpour