← Back to all news
arXiv Statistics ML October 1, 2026 By Zijun Chen, Zihan Zhang

Tighter Regret Bounds for Contextual Action-Set Reinforcement Learning

Read the original on arXiv Statistics ML →

The Flow has not summarised this story yet — read it at arXiv Statistics ML.

  • reinforcement-learning
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Jul 23

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence

arXiv:2607. 19854v1 Announce Type: new Abstract: We study horizon-free regret minimization for finite-horizon time-homogeneous tabular Markov decision processes with $S$ states, $A$ actions, horizon $H$, and per-trajectory total reward bounded by $1$.

By Runlong Zhou, Zihan Zhang, Maryam Fazel, Simon S. Du
reinforcement-learning
More like this →
arXiv Machine Learning
4d ago

Optimal Multi-Reward Reinforcement Learning

arXiv:2609.36486v1 Announce Type: new Abstract: We study an unknown-transition finite-horizon Markov decision process (MDP) with a finite collection of known reward functions $\{r^1, r^2, \ldots, r^M...

By Zijun Chen, Zihan Zhang
reinforcement-learning
More like this →
arXiv Machine Learning
Jun 2

Online Learning in MDPs with Partially Adversarial Transitions and Losses

arXiv:2602. 09474v2 Announce Type: replace Abstract: We study reinforcement learning in MDPs whose transition function is stochastic at most steps but may behave adversarially at a fixed subset of $\Lambda$ steps per episode.

By Ofir Schlisselberg, Tal Lancewicki, Yishay Mansour
reinforcement-learningsafety
More like this →
arXiv Machine Learning
Jun 25

Minimax PAC Bounds for Learning in Exogenous Contextual MDPs

arXiv:2606. 25170v1 Announce Type: cross Abstract: We study PAC learning in tabular discounted Markov decision processes with exogenous i.

By Corentin Pla, Hugo Richard, Marc Abeille, Vianney Perchet
agentsreinforcement-learning
More like this →
arXiv Machine Learning
Jun 2

An Improved Algorithm for Adversarial Linear Contextual Bandits via Reduction

arXiv:2508. 11931v3 Announce Type: replace Abstract: We present an oracle-efficient, near-optimal algorithm for linear contextual bandits with adversarial losses and stochastic action sets, only requiring a linear optimization oracle for the action sets in each round.

By Tim van Erven, Jack Mayo, Julia Olkhovskaya, Chen-Yu Wei
safety
More like this →
arXiv Machine Learning
Aug 27

Minimax Alternating Regret for the Experts Problem and Online Convex Optimization

arXiv:2608. 25182v1 Announce Type: cross Abstract: In this paper, we study alternating regret in online convex optimization (OCO), motivated by the success of alternating learning dynamics in two-player games.

By Mengxiao Zhang
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea