返回
TPN:Triple network algorithm for deep reinforcement learning
DOI:10.1016/j.neucom.2024.127755.png)
摘要
En 中文
The target net method has been the foundation of deep reinforcement learning since Deepmind first proposed it in 2015. Almost all the current popular reinforcement learning algorithms include target net. However, while the slowly updated target network improves the stability of the algorithm, it also reduces the performance of the algorithm. In this paper, the authors design a novel triple-network algorithm(TPN). TPN combines the temporal-difference(TD) algorithm and policy gradient(PG) theorem. Using three networks to estimate the state value( u ), action value ( q ) , and policy( r ). These networks have no primary or secondary distinction but are trained synchronously and influence each other. The author found that through this TPN architecture, the convergence and stability of the algorithm can be greatly improved without increasing the amount of calculation. Although it is only a basic framework at present. The calculation process of TPN is simple and easy to implement. Experiments prove that the convergence speed and stability of TPN in discrete cases are better than PPO.
Keyword:
TPN
Deep reinforcement learning
Target net method
期刊
IF:
6.5
论文数:
2.5W
被引数:
6.5W
机构
引用论文
Study of bi-directional buck-boost converter topologies for application in electrical vehicle motor drives应用于电动汽车电机驱动的双向buck-boost变换器拓扑研究
Fabrication of porous hollow γ-Al2O3 nanofibers by facile electrospinning and its application for water remediation静电纺丝法制备多孔中空 γ-Al2O3纳米纤维及其在水体修复中的应用
ICT for informal workers in Sub-Saharan Africa: Systematic review and analysis撒哈拉以南非洲非正规工人的信通技术: 系统回顾和分析
Novel 3D printing-based probe for impedance spectroscopic examination of oral mucosa: design and preliminary testing with phantom models基于3D打印的新型探头用于口腔黏膜阻抗谱检测:设计与基于仿模的初步测试
A process dissociation framework: Separating automatic from intentional uses of memory过程分离框架: 将自动与有意使用的内存分开

