← Back to all news
arXiv Machine Learning August 26, 2026 By Qizhen Jia, Keqin Liu

From Relaxed Indexability to Exact Indexability: A $t$-Step Approach for Partially Observable Restless Bandits

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Jun 11

Restless bandits with imperfect binary feedback: PCL-indexability analysis and computation

arXiv:2606. 11192v1 Announce Type: new Abstract: We study restless bandits with binary latent states and imperfect binary feedback, motivated by opportunistic spectrum access with sensing errors.

By Jos\'e Ni\~no-Mora
benchmarks
More like this →
arXiv Statistics ML
Sep 2

Model Predictive Control is almost Optimal for Heterogeneous Restless Multi-armed Bandits

arXiv:2511.08097v2 Announce Type: replace-cross Abstract: We consider a general infinite horizon Heterogeneous Restless multi-armed Bandit (RMAB). Heterogeneity is a fundamental problem for many real...

By Dheeraj Narasimha, Nicolas Gast
reinforcement-learning
More like this →
arXiv Machine Learning
Jul 23

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence

arXiv:2607. 19854v1 Announce Type: new Abstract: We study horizon-free regret minimization for finite-horizon time-homogeneous tabular Markov decision processes with $S$ states, $A$ actions, horizon $H$, and per-trajectory total reward bounded by $1$.

By Runlong Zhou, Zihan Zhang, Maryam Fazel, Simon S. Du
reinforcement-learning
More like this →
arXiv Machine Learning
Aug 18

Last-Iterate Analyses of FTRL with the 1/2-Tsallis Entropy in Stochastic Bandits

arXiv:2510. 22819v3 Announce Type: replace Abstract: The convergence analysis of online learning algorithms is central to machine learning theory, where the last-iterate convergence is particularly important, as it captures the learner's actual decisions and describes the evolution of the learning process over time.

By Jingxin Zhan, Yuze Han, Zhihua Zhang
safety
More like this →
arXiv Machine Learning
Jun 30

Optimal Regret for Single Index Bandits

arXiv:2605. 09454v2 Announce Type: replace-cross Abstract: We study the $\textit{single-index bandit}$ problem, where rewards depend on an unknown one-dimensional projection of high-dimensional contexts through an unknown reward function.

By Devdan Dey, Sujoy Bhore, Avishek Ghosh
reinforcement-learning
More like this →
arXiv Machine Learning
Aug 19

Toward the Optimal Regret-Instability Trade-off in Multi-Armed Bandits

arXiv:2608. 17841v1 Announce Type: cross Abstract: Multi-armed bandit algorithms are evaluated by regret, yet comparable regret can coexist with different allocations across independent runs.

By Kaifei Wang, Yinyu Ye, Han Zhong
reinforcement-learning
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea