arXiv Machine Learning By Takumi Shioda, Kohei Terashima, Tatsuo Nagai

Exact Action Values Are Not Enough: Rollout-Verified Reinforcement Fine-Tuning of a Reasoning Model for Multi-Zone VAV Control

Read the original on arXiv Machine Learning →

arXiv:2607. 27914v1 Announce Type: new Abstract: Multi-zone variable-air-volume control must balance thermal comfort, indoor air quality, and electricity use across several continuous actuators.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 20

Comparative Field Deployment of Reinforcement Learning and Model Predictive Control for Residential HVAC

arXiv:2510. 01475v2 Announce Type: replace-cross Abstract: Model Predictive Control (MPC) has demonstrated significant performance improvements over today's control methods for residential Heating, Ventilation, and Air Conditioning (HVAC), but deploying MPC often requires substantial engineering effort.

By Ozan Baris Mulayim, Elias N. Pergantis, Levi D. Reyes Premer, Bingqing Chen, Guannan Qu, Kevin J. Kircher, Mario Berg\'es
arXiv Machine Learning
Sep 14

Very Exciting: Zero-Shot Model Predictive Control of Buildings via Excitation-Based Generalized Transfer Learning Models

The paper presents a zero‑shot model predictive control (MPC) approach for buildings that uses excitation‑based generalized transfer learning models. By pretraining on purposefully probed operational data from multiple source buildings, the authors demonstrate that these models achieve superior control performance in 32 simulated target buildings, outperforming both an online linear MPC and a PI controller. This method eliminates the need for target‑specific data, reducing setup cost and facilitating broader deployment of energy‑efficient MPC in the building sector.

By Fabian Raisch, Felix Koch, Zack Xuereb Conti, Christoph Goebel, Benjamin Tischler
arXiv Machine Learning
Sep 25

Certified Predictive Value-of-Advice Gating for Cost-Aware Language-Model Guidance in Reinforcement Learning

The paper proposes a method for selectively querying language‑model advice in reinforcement learning by predicting the value of potential responses and only querying when the expected benefit outweighs the cost. It introduces a certified, response‑contingent metareasoning framework that guarantees near‑optimal advice usage under certain assumptions, and demonstrates that a calibrated controller with Qwen2.5 advisors can improve task performance while drastically reducing the number of advice calls on the BabyAI benchmark.

By Ibne Farabi Shihab, Md Najmus Swaqeeb, Abu Sa-Adat Mohamed Moon-Im Al Ahsan