The paper presents a zero‑shot model predictive control (MPC) approach for buildings that uses excitation‑based generalized transfer learning models. By pretraining on purposefully probed operational data from multiple source buildings, the authors demonstrate that these models achieve superior control performance in 32 simulated target buildings, outperforming both an online linear MPC and a PI controller. This method eliminates the need for target‑specific data, reducing setup cost and facilitating broader deployment of energy‑efficient MPC in the building sector.
By Fabian Raisch, Felix Koch, Zack Xuereb Conti, Christoph Goebel, Benjamin Tischler
arXiv:2501. 13703v2 Announce Type: replace-cross Abstract: Transfer Learning (TL) is an emerging field in modeling building thermal dynamics.
By Fabian Raisch, Thomas Krug, Christoph Goebel, Benjamin Tischler
Buildings account for roughly one-third of global energy consumption and CO$_2$ emissions. Optimizing indoor climate systems plays a critical role for urban climate mitigation aligned with UN Sustainable Development Goals 11 and 13.
arXiv:2601. 23225v2 Announce Type: replace-cross Abstract: Deep reinforcement learning (RL) is increasingly deployed in resource-constrained environments, yet go-to function approximators - multilayer perceptrons (MLPs) - are often parameter-inefficient due to an imperfect inductive bias for the smooth structure of many value functions.
By Rajib Mostakim, Reza T. Batley, Sourav Saha
arXiv:2608. 19804v1 Announce Type: new Abstract: Buildings account for roughly one-third of global energy consumption and CO$_2$ emissions.
By Xu Yang, Kailai Sun, Dianyu Zhong, Qianchuan Zhao
The paper introduces NOMAD‑RL, a reinforcement learning controller for HVAC systems that learns to adapt across diverse thermal zones via a universal thermostat interface. It employs an adaptive domain randomization scheme using physics‑informed normalizing flows to generate realistic, multimodal training data, enabling the recurrent policy to handle partial observability. Experiments show NOMAD‑RL outperforms constant‑setpoint PID and non‑randomized RL, and rivals well‑tuned model predictive control, especially in multi‑zone scenarios.
By Pablo Boitel, Kun Zhang