arXiv Machine Learning

Online Inference for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

arXiv:2608. 12973v1 Announce Type: cross Abstract: In this paper, we study how to perform statistical inference for quantile temporal difference learning (QTD) in distributional reinforcement learning.

arXiv AI
Sep 21

Deep Reinforcement Learning with Buffered Quantile Objectives

The paper introduces Deep-BQRL, a model‑free distributional reinforcement‑learning framework that extends buffered‑quantile learning to neural function approximation. It learns conditional return quantiles from sampled transitions, constructs buffered action scores, and uses ensemble disagreement for exploration, enabling risk‑sensitive decision‑making without explicit return‑law planning. Experiments on asset‑selling and slippery FrozenLake show that Deep‑BQRL achieves smaller mean cumulative point‑quantile policy gaps than PPO and TRPO, while illustrating interpretable risk‑sensitive stopping decisions.

By Mohammad Alipour-vaezi, Sajad Khodadadian
arXiv Machine Learning
Sep 10

Decision-Centered Abstractions via Orthogonal Estimation of Difference-of-Q Functions

The paper introduces state abstractions that preserve the difference of Q‑functions for offline reinforcement learning, aiming to exclude irrelevant dynamics from rich state data. It proposes a dynamic generalization of the R‑learner that uses orthogonal estimation and sparse learning to estimate the Q‑function contrast, achieving faster convergence and consistency under a margin condition. Experiments on simulated and simulator‑augmented real data show variance reductions and demonstrate that the necessary information for sequential decision‑making can be smaller than that required for full state prediction.

By Defu Cao, Angela Zhou
arXiv AI
Jun 11

Generalizing Beyond Suboptimality: Offline Reinforcement Learning Learns Effective Scheduling through Random Solutions

arXiv:2509. 10303v2 Announce Type: replace-cross Abstract: Online reinforcement learning (RL) approaches have demonstrated strong performance on Job Shop Scheduling (JSP) and Flexible JSP (FJSP) problems by learning scheduling policies through direct interaction with simulated environments.

By Jesse van Remmerden, Zaharah Bukhsh, Yingqian Zhang
arXiv Machine Learning
Jul 8

Model-based Bootstrap of Controlled Markov Chains

arXiv:2605. 12410v2 Announce Type: replace-cross Abstract: We propose and analyze a model-based bootstrap for transition kernels in finite controlled Markov chains (CMCs) with possibly nonstationary or history-dependent control policies, a setting that arises naturally in offline reinforcement learning (RL) when the behavior policy generating the data is unknown.

By Ziwei Su, Imon Banerjee, Diego Klabjan
arXiv AI
Jun 16

QPILOTS: Efficient Test-Time Q-Steering for Flow Policies

arXiv:2606. 14801v1 Announce Type: cross Abstract: Flow-matching and diffusion policies are expressive action generators, but optimizing them with temporal-difference reinforcement learning (RL) remains difficult.

By Yifan Ruan, Chenyang Cao, Andreas Burger, Ali Pesaranghader, Kaveh Kamali, Jaehong Kim, Nandita Vijaykumar, Alan Aspuru-Guzik, Igor Gilitschenski, Nicholas Rhinehart