1
Return

HVAC Control With Optimization-Guided Behavior Cloning and Restricted Residual Policy Learning

delete2026-02-16
delete0
PRE
AI
S
Shuhua Gao
H
Hailong WEN
F
Feng Wang
K
Kunmeng Yang
向程 cover
向程 (Cheng Xiang)
DOI:10.1109/TSG.2026.3665037delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
HVAC control can be formulated as an optimization problem balancing energy consumption and occupant comfort. Model-based approaches such as model predictive control (MPC) deliver strong performance when accurate thermal models are available but degrade under model mismatch and require costly system identification. Learning-based methods, notably reinforcement learning (RL), avoid explicit modeling yet often incurs unsafe exploration, poor initial performance, and slow convergence in practice. We propose Optimization-Guided Behavioral Cloning with Restricted Residual Policy Learning (OGBC+RRPL), a hybrid framework that unites model-based optimization and data-driven adaptation. OGBC first builds an optimization-based HVAC control expert by solving a linear program derived from an approximate autoregressive thermal model and trains an initial policy offline via behavioral cloning. RRPL then refines this policy online by learning a compensating lightweight residual policy with an actor-critic algorithm enhanced by action clipping and critic regularization to ensure stable and safe updates. This design preserves optimization knowledge, mitigates catastrophic forgetting, and constrains risky exploration during online adaptation. Experiments on CityLearn v2 demonstrate that OGBC+RRPL delivers high-quality control immediately upon deployment, attains comparable or better performance far faster than standard online RL algorithms, and outperforms offline RL methods on both energy consumption and temperature-violation metrics. The results indicate that infusing optimization guidance into residual online learning provides a practical, data-efficient, and safe approach for real-world HVAC control. Our implementation code can be found at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/HailongW0818/HVAC-OGBC-RRPL</uri>
Keywords:
HVAC control
behavior cloning
residual policy learning
reinforcement learning
optimization

Journal

IEEE Transactions on Smart Grid cover
IEEE Transactions on Smart Grid
IF:
9.8
Papers:
5.6K
Citations:
4.3W

Organization

S
shandong university
Scholars:
9.1W
Papers: 6.3W
Citations: 94
S
State Grid Shandong Electric Power Company
Scholars:
108
Papers: 51
Citations: 0
N
National University of Singapore
Scholars:
7.4W
Papers: 6.4W
Citations: 11.4W
Cited Papers

Cited Papers

Citing Papers

Citing Papers