arXiv AI By Alfred Harwood, Jose Faustino, Alex Altair

Calculating Mutual Information between a Reward Maximizer and its Environment

Read the original on arXiv AI →

arXiv:2602. 12963v2 Announce Type: replace Abstract: An important question in the field of AI is the extent to which successful behaviour requires an internal representation of the world.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 16

Learning Policy from a Single Trajectory in Average-Reward Markov Decision Process

arXiv:2606. 16729v1 Announce Type: new Abstract: While there is an extensive body of work characterizing the sample complexity of discounted cumulative-reward MDPs, finite sample analyses for average-reward MDPs have been limited, and most existing works rely on restrictive assumptions such as ergodicity or access to a generative model.

By Jongmin Lee, Ernest K. Ryu, Vaneet Aggarwal
arXiv Machine Learning
Sep 14

Independent Learning of Nash Equilibria in Partially Observable Markov Potential Games with Decoupled Dynamics

The paper investigates learning Nash equilibria in partially observable Markov games (POMGs) where agents cannot fully observe the state. By focusing on a subclass with independent state transitions and a Markov potential game structure, the authors propose an independent learning algorithm that allows agents to converge to an approximate Nash equilibrium using only their own observations and actions, without communication. Under a filter stability assumption, finite‑history policies are shown to approximate the POMG sufficiently, enabling a surrogate near‑potential Markov game and yielding quasi‑polynomial sample and computational complexity.

By Philip Jordan, Maryam Kamgarpour
arXiv Machine Learning
Sep 23

Tight Sample Complexity Bounds for Entropic Best Policy Identification

The paper investigates best‑policy identification in finite‑horizon, risk‑sensitive reinforcement learning using the entropic risk measure. It identifies a gap between known lower bounds ≥ η(e^{|eta|H}) and upper bounds ≤ O(e^{2|eta|H}) for sample complexity, attributing the excess factor to loose concentration bounds for exponential utilities. By employing a forward‑model algorithm with KL‑based exploration bonuses and a novel stopping rule, the authors achieve a sample complexity that matches the lower bound, closing the previously open exponential gap.

By Amer Essakine, Claire Vernade
arXiv Machine Learning
1d ago

When Do Intrinsic Rewards Lead to Exploration?

The paper investigates when intrinsic rewards effectively drive exploration in reinforcement learning. It introduces a formal criterion that evaluates policies based on the counterfactual information they acquire, comparing how well their histories can replace experience from alternative policies. Using a simple environment, the authors show that common intrinsic reward objectives—count-based, prediction-error, empowerment, and information-gain—can lead to Pareto-suboptimal exploration under this criterion, and they propose conditions and a new objective that better align with optimal exploration.

By Scott W. Viteri (Stanford University), Laura Gomezjurado Gonzalez (Stanford University), Clark Barrett (Stanford University)