arrow
Return

Optimizing scheduling policy in smart grids using probabilistic Delayed Double Deep Q-Learning (P3DQL) algorithm

delete2022-10-01
delete11
PRE
AI
H
Hossein Mohammadi Rouzbahani *
H
Hadis Karimipour
L
Lei Lei
DOI:10.1016/j.seta.2022.102712delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
High penetration of smart devices in IoE-enabled smart grids besides decentralization originated from employing renewable resources face the power system with intricate optimization problems. Operation scheduling of energy components is one of the principal problems that need to be addressed. However, engaging with big data produced by the interconnected infrastructures, besides the high dimensional and uncertain environment, make traditional methods incapable of addressing these problems since exact modeling of the environment under uncertainties is impracticable. While learning-based methods suffer from excessive complexity and the curse of dimensionality, Deep Reinforcement Learning has recently successfully handled highly complex scheduling problems. However, biases and model efficiency are two primary considerations that need more investigation. Positive and negative biases lead to better exploration and exploitation, respectively, and their harmony, considering model efficiency, results in a better outcome. Accordingly, this research develops and introduces a novel algorithm named Probabilistic Delayed Double Deep Q-Learning, which is a combination of the tuned version of Double Deep Q-Learning and Delayed Q-Learning. The proposed algorithm makes a trade-off between overestimation and underestimation biases, guaranteeing sample complexity by applying a delay in updating the rule. Finally, the proposed algorithm is tested on three real-world datasets assessing its performance in various benchmarks. The results indicate that the developed model is thoroughly stable since both population and characteristic stability indices are less than 0.1 in all case studies. The average model's error is 0.028 showing the exactitude of the model while running time is lower than other examined methods. Utilizing the developed algorithm results in 11.1 % reduction in the average power ratio. Consequently, the peak load decreased from 8.043 kW to 5.8137 kW, resulting in a 30.1 % cost reduction.
Keywords:
Energy scheduling
Smart grid
Deep reinforcement learning
Double Q-Learning
Probably approximately correct learning
Internet of energy

Journal

Sustainable Energy Technologies and Assessments cover
Sustainable Energy Technologies and Assessments
IF:
7
Papers:
4.7K
Citations:
2.2W

Organization

U
University of Calgary
Scholars:
3.8W
Papers: 3.3W
Citations: 52
U
University of Guelph
Scholars:
1.3W
Papers: 1.2W
Citations: 1.7W
Cited Papers

Cited Papers

errShare
errSave
Optimal Energy Supply Scheduling for a Single Household: Integrating Machine Learning for Power Forecasting
err2019-09-01
err6
PREAI
errBuechler, Thomas; Pagel, Fabian; Petitjean, Thibault; Draz, Mahmoud; Albayrak, Sahin
errShare
errSave
An Adversarial Perturbation Oriented Domain Adaptation Approach for Semantic Segmentation
err2020-04-03
err0
errOAAI
errJihan Yang; Ruijia Xu; Ruiyu Li; Xiaojuan Qi; Xiaoyong Shen; Guanbin Li; Liang Lin
errShare
errSave
errShare
errSave
On-Line Building Energy Optimization Using Deep Reinforcement Learning
err2019-07-01
err126
errOAAI
errMocanu, Elena; Mocanu, Decebal Constantin; Nguyen, Phuong H.; Liotta, Antonio; Webber, Michael E.; Gibescu, Madeleine; Slootweg, J. G.
errShare
errSave
researcher View more