← Back to all news
arXiv Machine Learning September 30, 2026 By Yige Hong, Xiangcheng Zhang, Qiaomin Xie, Yudong Chen, Weina Wang

Achieving an $O(1/N)$ Optimality Gap in Average-Reward Weakly-Coupled MDPs

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • reinforcement-learning

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Statistics ML
Sep 2

Model Predictive Control is almost Optimal for Heterogeneous Restless Multi-armed Bandits

arXiv:2511.08097v2 Announce Type: replace-cross Abstract: We consider a general infinite horizon Heterogeneous Restless multi-armed Bandit (RMAB). Heterogeneity is a fundamental problem for many real...

By Dheeraj Narasimha, Nicolas Gast
reinforcement-learning
More like this →
arXiv Machine Learning
Jun 15

Lyapunov-Based Sample Complexity Analysis for Weakly-Coupled MDPs

arXiv:2606. 14095v1 Announce Type: new Abstract: We study the sample complexity of learning in average-reward weakly-coupled Markov decision processes (WCMDPs) and Restless Bandits (RBs) under a generative model.

By Tianhao Wu, Matthew Zurek, Weina Wang, Qiaomin Xie
reinforcement-learningsafety
More like this →
arXiv Machine Learning
Jun 29

Learning in Markovian bandits with non-observable states and constrained decision epochs

arXiv:2606. 27448v1 Announce Type: new Abstract: This paper studies the problem of regret minimization in Markovian bandits with \emph{non-observable states} and possibly \emph{constrained} decision epochs.

By Thomas Hira, Victor Boone, Urtzi Ayesta, Ina Maria Verloop
reinforcement-learningbenchmarkssafety
More like this →
arXiv Machine Learning
Jul 23

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence

arXiv:2607. 19854v1 Announce Type: new Abstract: We study horizon-free regret minimization for finite-horizon time-homogeneous tabular Markov decision processes with $S$ states, $A$ actions, horizon $H$, and per-trajectory total reward bounded by $1$.

By Runlong Zhou, Zihan Zhang, Maryam Fazel, Simon S. Du
reinforcement-learning
More like this →
arXiv Machine Learning
Jul 30

Top-$k$ Pareto Bandits: Hypervolume Regret for Multi-Objective Slate Selection

arXiv:2607. 26273v1 Announce Type: new Abstract: We consider a stochastic multi-objective bandit problem where, at each round, the agent selects a slate of $k$ arms and observes their $d$-dimensional reward vectors under semi-bandit feedback.

By Nicolas Gutowski, Fabien Chhel, Alexandre Letard, Sylvain Lamprier
agentsreinforcement-learning
More like this →
arXiv Machine Learning
1d ago

Nonpreemptive Scheduling While Learning Context-Dependent Service Rates

arXiv:2609.37660v1 Announce Type: new Abstract: We study nonpreemptive contextual queueing bandits in a single-server system. Each job is represented by a $d$-dimensional context vector; in each roun...

By Wansoo Choi, Seoungbin Bae, Dabeen Lee
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea