返回
Optimizing scheduling policy in smart grids using probabilistic Delayed Double Deep Q-Learning (P3DQL) algorithm
DOI:10.1016/j.seta.2022.102712.png)
摘要
En 中文
High penetration of smart devices in IoE-enabled smart grids besides decentralization originated from employing renewable resources face the power system with intricate optimization problems. Operation scheduling of energy components is one of the principal problems that need to be addressed. However, engaging with big data produced by the interconnected infrastructures, besides the high dimensional and uncertain environment, make traditional methods incapable of addressing these problems since exact modeling of the environment under uncertainties is impracticable. While learning-based methods suffer from excessive complexity and the curse of dimensionality, Deep Reinforcement Learning has recently successfully handled highly complex scheduling problems. However, biases and model efficiency are two primary considerations that need more investigation. Positive and negative biases lead to better exploration and exploitation, respectively, and their harmony, considering model efficiency, results in a better outcome. Accordingly, this research develops and introduces a novel algorithm named Probabilistic Delayed Double Deep Q-Learning, which is a combination of the tuned version of Double Deep Q-Learning and Delayed Q-Learning. The proposed algorithm makes a trade-off between overestimation and underestimation biases, guaranteeing sample complexity by applying a delay in updating the rule. Finally, the proposed algorithm is tested on three real-world datasets assessing its performance in various benchmarks. The results indicate that the developed model is thoroughly stable since both population and characteristic stability indices are less than 0.1 in all case studies. The average model's error is 0.028 showing the exactitude of the model while running time is lower than other examined methods. Utilizing the developed algorithm results in 11.1 % reduction in the average power ratio. Consequently, the peak load decreased from 8.043 kW to 5.8137 kW, resulting in a 30.1 % cost reduction.
Keyword:
Energy scheduling
Smart grid
Deep reinforcement learning
Double Q-Learning
Probably approximately correct learning
Internet of energy
期刊
IF:
7
论文数:
4.5K
被引数:
2.2W
机构
引用论文
Modeling and Optimizing Energy Supply and Demand in Home Area Power Network (HAPN)家庭区域电网 (HAPN) 中的能源供需建模与优化
IEEE ACCESS
IF3.6

