arrow
Return

TPN:Triple network algorithm for deep reinforcement learning

delete2024-07-01
delete1
PRE
AI
H
Han Chen
X
Xuanyin Wang *
DOI:10.1016/j.neucom.2024.127755delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The target net method has been the foundation of deep reinforcement learning since Deepmind first proposed it in 2015. Almost all the current popular reinforcement learning algorithms include target net. However, while the slowly updated target network improves the stability of the algorithm, it also reduces the performance of the algorithm. In this paper, the authors design a novel triple-network algorithm(TPN). TPN combines the temporal-difference(TD) algorithm and policy gradient(PG) theorem. Using three networks to estimate the state value( u ), action value ( q ) , and policy( r ). These networks have no primary or secondary distinction but are trained synchronously and influence each other. The author found that through this TPN architecture, the convergence and stability of the algorithm can be greatly improved without increasing the amount of calculation. Although it is only a basic framework at present. The calculation process of TPN is simple and easy to implement. Experiments prove that the convergence speed and stability of TPN in discrete cases are better than PPO.
Keywords:
TPN
Deep reinforcement learning
Target net method

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

Z
zhejiang university
Scholars:
17.7W
Papers: 12.1W
Citations: 152
Cited Papers

Cited Papers

Model-based Reinforcement Learning: A Survey
err2023-01-01
err161
errOAAI
errMoerland, Thomas M.; Broekens, Joost; Plaat, Aske; Jonker, Catholijn M.
errShare
errSave
errShare
errSave
ICT for informal workers in Sub-Saharan Africa: Systematic review and analysis
err2017-09-01
err0
PREAI
errNasibu Mramba; Joel Rumanyika; Mikko Apiola; Jarkko Suhonen
errShare
errSave
errShare
errSave
An Improved DDPG and Its Application Based on the Double-Layer BP Neural Network
err2020-01-01
err30
errOAAI
errZhang, Mingli; Zhang, Yijie; Gao, Zhengjie; He, Xiaolong
errShare
errSave
researcher View more