arXiv:2607. 19854v1 Announce Type: new Abstract: We study horizon-free regret minimization for finite-horizon time-homogeneous tabular Markov decision processes with $S$ states, $A$ actions, horizon $H$, and per-trajectory total reward bounded by $1$.
By Runlong Zhou, Zihan Zhang, Maryam Fazel, Simon S. Du
arXiv:2608. 07301v1 Announce Type: new Abstract: We study what can be recovered about the transition probabilities of a Markov decision process from optimal actions alone.
By Neal Batra
arXiv:2608. 17841v1 Announce Type: cross Abstract: Multi-armed bandit algorithms are evaluated by regret, yet comparable regret can coexist with different allocations across independent runs.
By Kaifei Wang, Yinyu Ye, Han Zhong
The paper studies a stationary decentralized Markov game where a focal agent experiences drifting rewards and dynamics due to learning peers, framing this as an agent‑centric continual reinforcement‑learning problem. It introduces the concept of an invariant core—maximal abstract patterns common to many successful trajectories—and proves a worst‑case conditioning theorem linking trajectory‑law drift to success coverage. The authors provide theoretical guarantees for survival horizon, first‑exit law, and regret, and validate their predictions with solvable models and empirical studies in continual control, cue‑MNIST, and Level‑Based Foraging.
By Dane Malenfant
Multi-armed bandit algorithms are evaluated by regret, yet comparable regret can coexist with different allocations across independent runs. We study the trade-off between worst-case regret $\mathcal{R}_{K,T}$ and instability $\mathcal S_{K,T}$, defined as the largest standard deviation of a terminal pull count, for $K$ arms and $T$ rounds.
arXiv:2607. 16858v1 Announce Type: cross Abstract: Across environments with mixed sources of uncertainty, unsupervised reinforcement learning requires intrinsic motivation that does not precommit to a particular direction of surprise.
By Alireza Furutanpey, Schahram Dustdar
The paper introduces Anchored Neighborhood Optimization (ANO), a new policy‑optimization method that directly designs a smooth, bounded gain field for the probability‑ratio surrogate objective. ANO anchors the identity map at a ratio of one, peaks at a specified trust‑region boundary, and limits the influence of extreme off‑policy samples while providing a bounded, redescending pull on outliers. Empirical results show ANO consistently outperforms existing methods on Atari and MuJoCo benchmarks, and it remains robust under aggressive learning‑rate settings.
By Yiheng Zhang, Yiming Wang, Kaiyan Zhao, Zhenglin Wan, Jiayu Chen, Leong Hou U
arXiv:2607. 08925v1 Announce Type: new Abstract: Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training rather than trade them off against return, as constrained Markov decision process (MDP) formulations do.
By Elham Daneshmand, Majid Khadiv, Glen Berseth, Hsiu-Chin Lin
arXiv:2607. 11920v1 Announce Type: cross Abstract: Evaluating decisions made under uncertainty is hard when labeled outcomes are scarce, costly, or confounded with luck.
By Jeff Helzner
arXiv:2607. 17316v1 Announce Type: new Abstract: The softmax policy $\pi(a \mid s) \propto \exp(\beta Q(s,a))$ is the default model of stochastic choice in reinforcement learning (RL).
By Silviu Pitis
arXiv:2608. 12753v1 Announce Type: new Abstract: We study decentralized multi-player reinforcement learning in episodic tabular Markov decision processes (MDPs) under three forms of information asymmetry: (A) unobserved actions with common rewards, (B) observed actions with independent rewards, and (C) unobserved actions with independent rewards.
By Larissa Xu, King Bi, William Chang
arXiv:2607. 07538v1 Announce Type: new Abstract: Training a model with noisy gradient descent can be idealized as overdamped Langevin dynamics on the loss landscape, and a natural safety question is to bound the probability $\nu_t(\mathcal{A}_H) = \mathbb{P}(Q_t \in \mathcal{A}_H)$ that the trajectory lies in a designated failure region $\mathcal{A}_H$.
By Adam M. Oberman