arXiv:2608.28791v1 Announce Type: new
Abstract: Real-time decision-making for enhanced geothermal systems (EGS) is challenging because long-term production periods involve high-dimensional control sp...
By Ruimin Dai, Guodong Chen, Randy Harsuko, Kunpeng Liu, Nori Nakata
arXiv:2603. 09344v3 Announce Type: replace Abstract: Offline reinforcement learning (RL) enables data-efficient and safe policy learning without online exploration, but its performance often degrades under distribution shift.
By Hongqiang Lin, Zhenghui Fu, Weihao Tang, Pengfei Wang, Yiding Sun, Qixian Huang, Dongxu Zhang
arXiv:2606. 10228v1 Announce Type: cross Abstract: Safe exploration is a prerequisite for deploying reinforcement learning (RL) agents in safety-critical domains.
By Kaustubh Mani, Yann Pequignot, Vincent Mai, Liam Paull
The tutorial titled "Deep Learning for Sequential Decision Making under Uncertainty: Foundations, Frameworks, and Frontiers" explores how modern deep learning techniques—such as neural networks, transformers, large language models, and deep reinforcement learning—can be integrated with operations research and management science to address complex, uncertain, and dynamic decision problems. It argues that deep learning should complement, not replace, optimization, offering adaptability and scalable approximation while OR/MS provides rigorous constraint and uncertainty modeling. The tutorial organizes the field around predict‑then‑optimize, decision‑aware learning, constraint‑aware decision generation, and deep reinforcement learning, and highlights applications across supply chains, healthcare, energy, and autonomous systems.
By I. Esra Buyuktahtakin
The paper introduces BUMEX, a reinforcement learning exploration strategy that leverages a set of prior models containing the true transition kernel and reward function. By optimizing over this model set, the method derives upper and lower bounds on the Q‑function to guide exploration, providing theoretical guarantees of convergence to the optimal policy. When the model set follows a bounded‑parameter MDP structure, the optimization becomes convex, enabling finite‑time convergence under mild assumptions and demonstrating accelerated learning in simulations.
By J. S. van Hulst, W. P. M. H. Heemels, D. J. Antunes
arXiv:2606. 10705v1 Announce Type: cross Abstract: Reinforcement learning promises to optimize sequential decisions in large-scale systems.
By Yavar Yeganeh, Mahsa Shekari, Nicla Frigerio, Daniele Pagano, Andrea Matta