arXiv AI

Deep Reinforcement Learning with Buffered Quantile Objectives

The paper introduces Deep-BQRL, a model‑free distributional reinforcement‑learning framework that extends buffered‑quantile learning to neural function approximation. It learns conditional return quantiles from sampled transitions, constructs buffered action scores, and uses ensemble disagreement for exploration, enabling risk‑sensitive decision‑making without explicit return‑law planning. Experiments on asset‑selling and slippery FrozenLake show that Deep‑BQRL achieves smaller mean cumulative point‑quantile policy gaps than PPO and TRPO, while illustrating interpretable risk‑sensitive stopping decisions.

arXiv AI
Jun 11

Generalizing Beyond Suboptimality: Offline Reinforcement Learning Learns Effective Scheduling through Random Solutions

arXiv:2509. 10303v2 Announce Type: replace-cross Abstract: Online reinforcement learning (RL) approaches have demonstrated strong performance on Job Shop Scheduling (JSP) and Flexible JSP (FJSP) problems by learning scheduling policies through direct interaction with simulated environments.

By Jesse van Remmerden, Zaharah Bukhsh, Yingqian Zhang
arXiv Machine Learning
Jun 10

Discovering Interpretable Multi-Parameter Control Policies for Evolutionary Algorithms Using Deep Reinforcement Learning

arXiv:2606. 10129v1 Announce Type: new Abstract: While deep Reinforcement Learning (deep-RL) has been increasingly applied to parameter control in evolutionary algorithms, rigorous theoretical analysis of parameter control remains largely restricted to single-parameter settings, owing to the difficulty of deriving effective, interpretable multi-parameter policies amenable to formal study.

By Tai Nguyen, Phong Le, Carola Doerr, Nguyen Dang
Hugging Face Trending Papers
5d ago

Q-learning Penalized Transformer for Safe Offline Reinforcement Learning

The paper introduces Q-learning Penalized Transformer (QPT), a training–inference consistent framework for safe offline reinforcement learning. QPT trains a Transformer policy that generates actions conditioned on trajectory context and target return/cost while incorporating a Q-shaped penalty to balance safety, reward maximization, and behavior regularization. The method consistently outperforms strong baselines on 38 DSRL benchmark tasks and adapts robustly to varying constraint thresholds.