The paper introduces NOMAD‑RL, a reinforcement learning controller for HVAC systems that learns to adapt across diverse thermal zones via a universal thermostat interface. It employs an adaptive domain randomization scheme using physics‑informed normalizing flows to generate realistic, multimodal training data, enabling the recurrent policy to handle partial observability. Experiments show NOMAD‑RL outperforms constant‑setpoint PID and non‑randomized RL, and rivals well‑tuned model predictive control, especially in multi‑zone scenarios.
By Pablo Boitel, Kun Zhang
The paper presents a zero‑shot model predictive control (MPC) approach for buildings that uses excitation‑based generalized transfer learning models. By pretraining on purposefully probed operational data from multiple source buildings, the authors demonstrate that these models achieve superior control performance in 32 simulated target buildings, outperforming both an online linear MPC and a PI controller. This method eliminates the need for target‑specific data, reducing setup cost and facilitating broader deployment of energy‑efficient MPC in the building sector.
By Fabian Raisch, Felix Koch, Zack Xuereb Conti, Christoph Goebel, Benjamin Tischler
arXiv:2609.21108v1 Announce Type: new
Abstract: Deep reinforcement learning (DRL) has achieved strong performance across a wide range of continuous-control problems. These continuous-control policies...
By Sachini Weerasekara, Sagar Kamarthi, Jacqueline Isaacs
arXiv:2409. 19716v2 Announce Type: replace-cross Abstract: Constrained Reinforcement Learning (RL) has emerged as a significant research area within RL, where integrating constraints with rewards is crucial for enhancing safety and performance across diverse control tasks.
By Baohe Zhang, Lilli Frison, Thomas Brox, Joschka B\"odecker
arXiv:2606. 02049v1 Announce Type: new Abstract: The increasing integration of renewable energy sources into power systems, particularly in buildings equipped with photovoltaic (PV) panels and energy storage systems, introduces significant complexity in energy systems.
By Hallah Shahid Butt, Qiong Huang, G\"okhan Demirel, Kevin F\"orderer, Erfan Tajalli-Ardekani, Simnon Waczowicz, Luigi Spatafora, Veit Hagenmeyer, Benjamin Sch\"afer
arXiv:2608. 10634v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL), which learns environment dynamics to generate synthetic experience, is a promising approach to sample-efficient decision making.
By Zefeng Liang, Jie Qiao, Ruichu Cai, Weilin Chen, Zhifeng Hao
BVR Sim is an open‑source, Gymnasium‑style environment for heterogeneous air‑combat reinforcement learning, supporting multiple JSBSim aircraft models (F‑15, F‑16, F/A‑18, F‑22) with configurable weapons, sensors, and opponents. It offers a unified tactical action interface, interchangeable Python and accelerated C++ backends, entity‑oriented observations, compositional rewards, scripted opponents, replay and visualization, and adapters for multi‑agent learning frameworks. At a 0.4‑second decision interval, the C++ backend achieves 104 simulated seconds per wall‑clock second in 1‑vs‑1 and remains practical through 10‑vs‑10 scenarios, and a policy trained on the F‑16 transfers to four unseen aircraft with a 45.5% mean win rate after controller adaptation.
By Haocheng Sun (Beijing University of Posts,Telecommunications), Mulai Tan (Air Force Engineering University)
Buildings account for roughly one-third of global energy consumption and CO$_2$ emissions. Optimizing indoor climate systems plays a critical role for urban climate mitigation aligned with UN Sustainable Development Goals 11 and 13.
arXiv:2602. 12643v2 Announce Type: replace-cross Abstract: We present Unified Latent Dynamics (ULD), a novel reinforcement learning algorithm that unifies the efficiency of model-free methods with the representational strengths of model-based approaches, without incurring planning overhead.
By Jashaswimalya Acharjee, Balaraman Ravindran
arXiv:2509. 11259v2 Announce Type: replace-cross Abstract: Recent advancements in machine learning have largely been driven by foundation models (FMs) trained on large, diverse datasets, enabling them to generalize effectively to new, related tasks.
By David Schiff, Ofir Lindenbaum, Yonathan Efroni
arXiv:2304.10041v2 Announce Type: replace
Abstract: This work investigates formal policy synthesis for continuous-state stochastic dynamic systems subject to high-level specifications expressed in li...
By Lening Li, Zhentian Qian, Jianan Xia, Qiren Geng, Huasheng Zhang, Liang Hu, Qishuang Li, Junqiang Lou
arXiv:2608.29061v1 Announce Type: new
Abstract: Offline goal-conditioned reinforcement learning (GCRL) aims to learn policies for reaching diverse goals entirely from fixed trajectory data. Long-hori...
By Soohyun Choi, Seonvin Cho, Songnam Hong