返回
Guided deterministic policy optimization with gradient-free policy parameters information
DOI:10.1016/j.eswa.2023.120693.png)
摘要
En 中文
Deep Deterministic Policy Gradient (DDPG) and Twin Delayed Deep Deterministic Policy Gradient (TD3) are two classical deterministic policy gradient algorithms. It is worth noting that the policies of both DDPG and TD3 are completely dependent on the gradient of critics. This will cause the policy to be unstable and easy to converge to the local optimum in the learning process. Although the idea of maximum entropy learning can provide more effective exploration, it can only be applied to the algorithm using stochastic policy, not to DDPG and TD3. In this paper, we propose a deterministic policy optimization method combining gradient-free policy parameters information (GFPPI). Specifically, we obtain a new set of policies by injecting Gaussian noise into the policy parameters, and then weight these policy parameters based on critics to obtain GFPPI. Finally, GFPPI is used as the regularization term of the policy optimization function to guide the policy update. GFPPI can mitigate premature policy convergence and facilitate exploration with optimistic principles. We provide the theoretical guarantee for monotonic improvement of expected cumulative return using augmented loss function with GFPPI, experimentally analyze the role of GFPPI in policy optimization and combine it with deterministic policy gradient information for policy optimization. The experiments on OpenAI gym demonstrate that GFPPI can improve sample efficiency and enable the algorithm to get higher performance.
Keyword:
Deterministic policy gradient
Premature convergence
Local optimum
Policy optimization
Exploration
Sample efficiency
期刊
IF:
7.5
论文数:
2.9W
被引数:
10.2W
机构
引用论文
Echocardiographic assessment of abnormal left ventricular relaxation in man.超声心动图评估男性异常左心室舒张。
Heart
IF0
Effect of Palm Kernel Meal as Melamine Urea Formaldehyde Adhesive Extender for Plywood Application: Using a Fourier Transform Infrared Spectroscopy (FTIR) Study棕榈仁粉作为三聚氰胺尿素甲醛粘合剂在胶合板应用中的效果: 使用傅里叶变换红外光谱 (FTIR) 研究
Purification and characterization of oxidoreductases-catalyzing carbonyl reduction of the tobacco-specific nitrosamine 4-methylnitrosamino-1-(3-pyridyl)-1-butanone (NNK) in human liver cytosol人肝细胞溶胶中烟草特异性亚硝胺4-甲基亚硝基氨基-1-(3-吡啶基)-1-丁酮 (NNK) 的氧化还原酶催化羰基还原的纯化和表征
Xenobiotica
IF0
没有更多内容

