arXiv:2607. 16204v1 Announce Type: new Abstract: Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environments.
By Darshan Deshpande
arXiv:2609.14327v1 Announce Type: new
Abstract: Variance penalization is a principled approach to risk-sensitive reinforcement learning (RL) that explicitly trades expected return for policy stabilit...
By Saunak Kumar Panda, Tong Li, Yisha Xiang, Ruiqi Liu
arXiv:2603. 04818v3 Announce Type: replace Abstract: Disruptions at critical logistics nodes pose severe risks to global supply chains, yet existing risk prediction systems typically prioritize forecasting accuracy without providing operationally interpretable early warnings.
By Zhiming Xue, Yujue Wang, Menghao Huo
The paper presents a deep reinforcement learning framework that explicitly incorporates uncertainty quantification for risk and anomaly identification in distribution network operation. It combines distributional and Bayesian DRL to separate total uncertainty into aleatoric (inherent risk) and epistemic (out‑of‑distribution anomalies) components. The epistemic estimates guide exploration during training and enable anomaly detection with fallback control during deployment, while aleatoric estimates assess intrinsic operational risk.
By Ziqi Zhang
arXiv:2602. 04879v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a cornerstone for fine-tuning Large Language Models (LLMs), with Proximal Policy Optimization (PPO) serving as the de facto standard algorithm.
By Penghui Qi, Xiangxin Zhou, Zichen Liu, Tianyu Pang, Chao Du, Min Lin, Wee Sun Lee
arXiv:2606. 10228v1 Announce Type: cross Abstract: Safe exploration is a prerequisite for deploying reinforcement learning (RL) agents in safety-critical domains.
By Kaustubh Mani, Yann Pequignot, Vincent Mai, Liam Paull