arXiv:2609. 14959v1 Announce Type: new Abstract: We study decentralized learning of Nash equilibria (NE) in infinite-horizon discounted Markov games under bandit feedback, focusing on Markov $\alpha$-potential games.
By S. Rasoul Etesami
The paper introduces adaptive policy portfolios for robust Markov decision processes, where a finite set of memoryless randomized policies is synthesized offline and paired with an online selector. It defines robust regret as a measure of portfolio quality, comparing each portfolio member to the optimal policy for each plausible environment. The authors provide a complexity-theoretic analysis of portfolio certification and synthesis, showing that even deterministic portfolios in simple settings are highly complex, and present an offline construction method that can be specialized at runtime.
By Kasper Engelen, Sebastian Junges, Guillermo A. P\'{e}rez, Marnix Suilen
The paper investigates adaptive policy portfolios for Robust Markov Decision Processes (RMDPs), proposing finite sets of memoryless randomized policies generated offline and selected online. It introduces robust regret as a metric for portfolio quality, comparing each portfolio member’s performance to the optimal policy for each plausible environment. The authors provide complexity-theoretic results showing that certifying and synthesizing such portfolios is highly intractable, and they present an offline construction method that can be specialized at runtime.
arXiv:2606. 27448v1 Announce Type: new Abstract: This paper studies the problem of regret minimization in Markovian bandits with \emph{non-observable states} and possibly \emph{constrained} decision epochs.
By Thomas Hira, Victor Boone, Urtzi Ayesta, Ina Maria Verloop
arXiv:2511.08097v2 Announce Type: replace-cross
Abstract: We consider a general infinite horizon Heterogeneous Restless multi-armed Bandit (RMAB). Heterogeneity is a fundamental problem for many real...
By Dheeraj Narasimha, Nicolas Gast
arXiv:2608. 06520v1 Announce Type: new Abstract: We study online cooperative control of a multi-agent system under Byzantine attacks.
By Ximing Sun, Yue Wang