arrow
返回

Inference-Based Posteriori Parameter Distribution Optimization

delete2022-05-01
delete7
PRE
AI
X
Xuesong Wang
T
Tianyi Li
Y
Yuhu Cheng *
陈晨 封面图
陈晨 (C. L. Philip Chen)
DOI:10.1109/TCYB.2020.3023127delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Encouraging the agent to explore has always been an important and challenging topic in the field of reinforcement learning (RL). Distributional representation for network parameters or value functions is usually an effective way to improve the exploration ability of the RL agent. However, directly changing the representation form of network parameters from fixed values to function distributions may cause algorithm instability and low learning inefficiency. Therefore, to accelerate and stabilize parameter distribution learning, a novel inference-based posteriori parameter distribution optimization (IPPDO) algorithm is proposed. From the perspective of solving the evidence lower bound of probability, we, respectively, design the objective functions for continuous-action and discrete-action tasks of parameter distribution optimization based on inference. In order to alleviate the overestimation of the value function, we use multiple neural networks to estimate value functions with Retrace, and the smaller estimate participates in the network parameter update; thus, the network parameter distribution can be learned. After that, we design a method used for sampling weight from network parameter distribution by adding an activation function to the standard deviation of parameter distribution, which achieves the adaptive adjustment between fixed values and distribution. Furthermore, this IPPDO is a deep RL (DRL) algorithm based on off-policy, which means that it can effectively improve data efficiency by using off-policy techniques such as experience replay. We compare IPPDO with other prevailing DRL algorithms on the OpenAI Gym and MuJoCo platforms. Experiments on both continuous-action and discrete-action tasks indicate that IPPDO can explore more in the action space, get higher rewards faster, and ensure algorithm stability.
Keyword:
Task analysis
Optimization
Training
Artificial neural networks
Linear programming
Inference algorithms
Markov processes
Exploration
inference
parameter distribution
reinforcement learning (RL)
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Cybernetics 封面图
IEEE Transactions on Cybernetics
IF:
10.5
论文数:
1.1W
被引数:
5.0W

机构

U
University of Macau
学者数:
1.1W
论文数: 1.3W
被引数: 2.0W
引用论文

引用论文

Seismic Upgradation of RC Beams Strengthened with Externally Bonded Spent Catalyst Based Ferrocement Laminates
err2023-03-01
err0
errOAAI
errR. Balamuralikrishnan; A. S. H. Al-Mawaali; M. Y. Y. Al-Yaarubi; B. B. Al-Mukhaini; Asima Kaleem
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
Proximal Parameter Distribution Optimization
err2021-06-01
err9
PREAI
errWang, Xuesong; Li, Tianyi; Cheng, Yuhu
err分享
err收藏
err分享
err收藏
err分享
err收藏
Composite Learning Enhanced Robot Impedance Control
err2020-03-01
err63
PREAI
errSun, Tairen; Peng, Liang; Cheng, Long; Hou, Zeng-Guang; Pan, Yongping
err分享
err收藏
学者 查看更多内容