Return
HVAC Control With Optimization-Guided Behavior Cloning and Restricted Residual Policy Learning
S
H
F
K
DOI:10.1109/TSG.2026.3665037.png)
Abstract
En 中文
HVAC control can be formulated as an optimization problem balancing energy consumption and occupant comfort. Model-based approaches such as model predictive control (MPC) deliver strong performance when accurate thermal models are available but degrade under model mismatch and require costly system identification. Learning-based methods, notably reinforcement learning (RL), avoid explicit modeling yet often incurs unsafe exploration, poor initial performance, and slow convergence in practice. We propose Optimization-Guided Behavioral Cloning with Restricted Residual Policy Learning (OGBC+RRPL), a hybrid framework that unites model-based optimization and data-driven adaptation. OGBC first builds an optimization-based HVAC control expert by solving a linear program derived from an approximate autoregressive thermal model and trains an initial policy offline via behavioral cloning. RRPL then refines this policy online by learning a compensating lightweight residual policy with an actor-critic algorithm enhanced by action clipping and critic regularization to ensure stable and safe updates. This design preserves optimization knowledge, mitigates catastrophic forgetting, and constrains risky exploration during online adaptation. Experiments on CityLearn v2 demonstrate that OGBC+RRPL delivers high-quality control immediately upon deployment, attains comparable or better performance far faster than standard online RL algorithms, and outperforms offline RL methods on both energy consumption and temperature-violation metrics. The results indicate that infusing optimization guidance into residual online learning provides a practical, data-efficient, and safe approach for real-world HVAC control. Our implementation code can be found at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/HailongW0818/HVAC-OGBC-RRPL</uri>
Keywords:
HVAC control
behavior cloning
residual policy learning
reinforcement learning
optimization
Journal
IF:
9.8
Papers:
5.6K
Citations:
4.3W
