← Back to all news
arXiv Machine Learning September 30, 2026 By Zijun Chen, Zihan Zhang

Optimal Multi-Reward Reinforcement Learning

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • reinforcement-learning

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Jul 23

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence

arXiv:2607. 19854v1 Announce Type: new Abstract: We study horizon-free regret minimization for finite-horizon time-homogeneous tabular Markov decision processes with $S$ states, $A$ actions, horizon $H$, and per-trajectory total reward bounded by $1$.

By Runlong Zhou, Zihan Zhang, Maryam Fazel, Simon S. Du
reinforcement-learning
More like this →
arXiv Machine Learning
Jun 25

Minimax PAC Bounds for Learning in Exogenous Contextual MDPs

arXiv:2606. 25170v1 Announce Type: cross Abstract: We study PAC learning in tabular discounted Markov decision processes with exogenous i.

By Corentin Pla, Hugo Richard, Marc Abeille, Vianney Perchet
agentsreinforcement-learning
More like this →
arXiv Machine Learning
Aug 18

Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs

arXiv:2509. 16586v2 Announce Type: replace Abstract: Recent advances have significantly improved our understanding of the sample complexity of learning in average-reward Markov decision processes (AMDPs) under the generative model.

By Yukuan Wei, Xudong Li, Lin F. Yang
reinforcement-learningsafety
More like this →
arXiv Machine Learning
Jun 16

Fast Non-Episodic Finite-Horizon RL with K-Step Lookahead Thresholding

arXiv:2602. 00781v2 Announce Type: replace Abstract: Online reinforcement learning in non-episodic, finite-horizon MDPs remains underexplored and is challenged by the need to estimate returns to a fixed terminal time.

By Jiamin Xu, Kyra Gan
reinforcement-learningbenchmarks
More like this →
arXiv Machine Learning
Jun 2

Online Learning in MDPs with Partially Adversarial Transitions and Losses

arXiv:2602. 09474v2 Announce Type: replace Abstract: We study reinforcement learning in MDPs whose transition function is stochastic at most steps but may behave adversarially at a fixed subset of $\Lambda$ steps per episode.

By Ofir Schlisselberg, Tal Lancewicki, Yishay Mansour
reinforcement-learningsafety
More like this →
arXiv AI
Jun 4

Exact Unlearning in Reinforcement Learning

arXiv:2606. 04182v1 Announce Type: cross Abstract: We formulate the problem of \emph{exact unlearning} in reinforcement learning, where the goal is to design an efficient framework that enables the removal of any user's data upon deletion request, i.

By Thanh Nguyen-Tang, Raman Arora
reinforcement-learning
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea