返回
Dynamic sparse coding-based value estimation network for deep reinforcement learning
DOI:10.1016/j.neunet.2023.09.013.png)
摘要
En 中文
Deep Reinforcement Learning (DRL) is one powerful tool for varied control automation problems. Performances of DRL highly depend on the accuracy of value estimation for states from environments. However, the Value Estimation Network (VEN) in DRL can be easily influenced by the phenomenon of catastrophic interference from environments and training. In this paper, we propose a Dynamic Sparse Coding-based (DSC) VEN model to obtain precise sparse representations for accurate value prediction and sparse parameters for efficient training, which is not only applicable in Q-learning structured discrete-action DRL but also in actor-critic structured continuous-action DRL. In detail, to alleviate interference in VEN, we propose to employ DSC to learn sparse representations for accurate value estimation with dynamic gradients beyond the conventional l1 norm that provides same-value gradients. To avoid influences from redundant parameters, we employ DSC to prune weights with dynamic thresholds more efficiently than static thresholds like l1 norm. Experiments demonstrate that the proposed algorithms with dynamic sparse coding can obtain higher control performances than existing benchmark DRL algorithms in both discrete-action and continuous-action environments, e.g., over 25% increase in Puddle World and about 10% increase in Hopper. Moreover, the proposed algorithm can reach convergence efficiently with fewer episodes in different environments.(c) 2023 Elsevier Ltd. All rights reserved.
Keyword:
Deep reinforcement learning
Value estimation network
Dynamic sparse coding
期刊
IF:
6.3
论文数:
8.2K
被引数:
3.0W
机构
引用论文
Safe reinforcement learning for real-time automatic control in a smart energy-hub智能能源中心实时自动控制的安全强化学习
APPLIED ENERGY
IF11
Deep reinforcement learning guided graph neural networks for brain network analysis用于脑网络分析的深度强化学习引导图神经网络
NEURAL NETWORKS
IF6.3
Discovering diverse solutions in deep reinforcement learning by maximizing state-action-based mutual information
NEURAL NETWORKS
IF6.3
A process dissociation framework: Separating automatic from intentional uses of memory过程分离框架: 将自动与有意使用的内存分开

