arrow
返回

A deep deterministic policy gradient algorithm based on averaged state-action estimation

delete2022-07-01
delete10
PRE
AI
X
Xu Jian
H
Haifei Zhang
J
Jianlin Qiu *
DOI:10.1016/j.compeleceng.2022.108015delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Deep Reinforcement Learning (DRL), one of the most popular research topics in artificial intelligence, has achieved a breakthrough in continuous control tasks. Nonetheless, the DRL algorithm's instability and local optimality have a bad influence impact on its performance. The Deep Deterministic Policy Gradients (DDPG) algorithm uses a soft update to slow down the target value rate of change to alleviate this problem. However, there is still a specific target approximate error variance. The variance will aggravate the degree of the data dispersion and reduce the stability of the model. This paper proposed the DDPG with averaged state-action estimation (Averaged-DDPG) algorithm. It aims to minimize the adverse effects of conflict, which calculates the action reward by averaging the estimated values of previously learned Q values, thus reducing the training process's fluctuation and improving the algorithm's performance. The evaluation results in continuous control tasks show that Averaged-DDPG can enhance the agent's learning efficiency and training balance more effectively than the original DDPG algorithm.
Keyword:
Deep reinforcement learning
Deep deterministic policy gradients
Averaged state-action estimation
Target approximate error

期刊

C
Computers and Electrical Engineering
IF:
4.9
论文数:
6.7K
被引数:
1.3W

机构

N
Nantong University
学者数:
1.9W
论文数: 1.1W
被引数: 2.0W
引用论文

引用论文

Predicting Water Quality Based On EEMD And LSTM Networks
err2021-05-22
err0
PREAI
errDingyuan Zhang; Renkai Chang; Haisheng Wang; Yong Wang; Hao Wang; Shaoqing Chen
err分享
err收藏
err分享
err收藏
Cancer care and research in India: what does it mean to Nepal?
err2014-07-01
err0
PREAI
errBishal Gyawali; Bishesh Poudyal; Tomoya Shimokata; Yuichi Ando
err分享
err收藏
Accelerating deep reinforcement learning model for game strategy
err2020-09-01
err17
PREAI
errLi, Yifan; Fang, Yuchun; Akhtar, Zahid
err分享
err收藏
Forming a new small sample deep learning model to predict total organic carbon content by combining unsupervised learning with semisupervised learning
err2019-10-01
err63
PREAI
errZhu, Linqi; Zhang, Chong; Zhang, Chaomo; Zhang, Zhansong; Nie, Xin; Zhou, Xueqing; Liu, Weinan; Wang, Xiu
err分享
err收藏
学者 查看更多内容