Hugging Face Trending Papers

The Curious Case of Exploding DecPOMDPs: Containing the Fire through Policy Counting

The paper addresses the exponential complexity of Decentralised Partially Observable Markov Decision Processes (DecPOMDPs) by shifting focus from counting agents to counting policies. By exploiting symmetry among agents, it introduces a compact encoding that reduces model complexity and evaluation cost to polynomial dependence. The authors further develop a policy‑counted dynamic programming algorithm that efficiently solves these policy‑counted DecPOMDPs.

arXiv AI
Aug 19

The Curious Case of Exploding DecPOMDPs: Containing the Fire through Policy Counting

The paper introduces policy‑counted DecPOMDPs, a variant of decentralized partially observable Markov decision processes that mitigates the exponential growth in agent numbers by counting policies instead of agents. By exploiting symmetry among agents, the authors achieve a compact representation that reduces model complexity and evaluation cost to polynomial levels. They further present a dynamic programming algorithm that leverages this compact form to solve policy‑counted DecPOMDPs efficiently.

By Nazl{\i} Nur Karabulut, tanya Braun
arXiv Machine Learning
2d ago

Policy Iteration Is Not Strongly Polynomial for Deterministic Markov Decision Processes: The Price of Algorithmic Anarchy

The paper presents an exponential lower bound on the number of iterations required by Howard's policy iteration algorithm for deterministic discounted Markov decision processes with at most two actions per state, when the discount factor is part of the input. This result shows that Howard's method cannot be strongly polynomial in this setting and establishes an exponential gap compared to the simplex method with Dantzig's pivoting rule, which remains strongly polynomial. Even with rewards limited to logarithmic bit length, a stretched‑exponential lower bound is achieved, highlighting a fundamental difference between decentralized, simultaneous improvements and Dantzig's coordinated single‑action selection.

By Han Zhong, Yinyu Ye
Hugging Face Trending Papers
Aug 18

Adaptive Policy Portfolios for Robust Markov Decision Processes

The paper investigates adaptive policy portfolios for Robust Markov Decision Processes (RMDPs), proposing finite sets of memoryless randomized policies generated offline and selected online. It introduces robust regret as a metric for portfolio quality, comparing each portfolio member’s performance to the optimal policy for each plausible environment. The authors provide complexity-theoretic results showing that certifying and synthesizing such portfolios is highly intractable, and they present an offline construction method that can be specialized at runtime.

arXiv AI
Aug 19

Adaptive Policy Portfolios for Robust Markov Decision Processes

The paper introduces adaptive policy portfolios for robust Markov decision processes, where a finite set of memoryless randomized policies is synthesized offline and paired with an online selector. It defines robust regret as a measure of portfolio quality, comparing each portfolio member to the optimal policy for each plausible environment. The authors provide a complexity-theoretic analysis of portfolio certification and synthesis, showing that even deterministic portfolios in simple settings are highly complex, and present an offline construction method that can be specialized at runtime.

By Kasper Engelen, Sebastian Junges, Guillermo A. P\'{e}rez, Marnix Suilen
arXiv Machine Learning
1d ago

Linear Programming Representations and Strongly Polynomial Algorithms for Robust Markov Decision Processes

The paper presents linear programming formulations and strongly polynomial algorithms for robust Markov decision processes (RMDPs) with rational polyhedral state-action rectangular uncertainty in rewards and transitions. By encoding a finite sequence of robust policy-iteration steps, a single LP is constructed whose optimal solutions recover the robust optimal value and all optimal stationary randomized policies. The authors provide a general complexity analysis of robust policy iteration, improving known bounds for α1 and α1∞ RMDPs and establishing new strongly polynomial bounds for general interval, weighted α1, and Wasserstein RMDPs, as well as turn‑based stochastic games with these uncertainty sets.

By Han Zhong, Yinyu Ye
arXiv Machine Learning
Sep 14

Independent Learning of Nash Equilibria in Partially Observable Markov Potential Games with Decoupled Dynamics

The paper investigates learning Nash equilibria in partially observable Markov games (POMGs) where agents cannot fully observe the state. By focusing on a subclass with independent state transitions and a Markov potential game structure, the authors propose an independent learning algorithm that allows agents to converge to an approximate Nash equilibrium using only their own observations and actions, without communication. Under a filter stability assumption, finite‑history policies are shown to approximate the POMG sufficiently, enabling a surrogate near‑potential Markov game and yielding quasi‑polynomial sample and computational complexity.

By Philip Jordan, Maryam Kamgarpour