返回
Stochastic Double Deep Q-Network
DOI:10.1109/ACCESS.2019.2922706.png)
摘要
En 中文
Estimation bias seriously affects the performance of reinforcement learning algorithms. The maximum operation may result in overestimation, while the double estimator operation often leads to underestimation. To eliminate the estimation bias, these two operations are combined together in our proposed algorithm named stochastic double deep Q-learning network (SDDQN), which is based on the idea of random selection. A tabular version of SDDQN is also given, named stochastic double Q-learning (SDQ). Both the SDDQN and SDQ are based on the double estimator framework. At each step, we choose to use either the maximum operation or the double estimator operation with a certain probability, which is determined by a random selection parameter. The theoretical analysis shows that there indeed exists a proper random selection parameter that makes SDDQN and SDQ unbiased. The experiments on Grid World and Atari 2600 games illustrate that our proposed algorithms can balance the estimation bias effectively and improve performance.
Keyword:
Estimation bias
deep reinforcement learning
maximum operation
double estimator operation
stochastic combination
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
暂无机构信息
引用论文
Study of bi-directional buck-boost converter topologies for application in electrical vehicle motor drives应用于电动汽车电机驱动的双向buck-boost变换器拓扑研究
Fabrication of porous hollow γ-Al2O3 nanofibers by facile electrospinning and its application for water remediation静电纺丝法制备多孔中空 γ-Al2O3纳米纤维及其在水体修复中的应用
Actor-Critic Off-Policy Learning for Optimal Control of Multiple-Model Discrete-Time Systems用于多模型离散时间系统最优控制的Actor-Critic Off-Policy学习
First-Principle Protocol for Calculating Ionization Energies and Redox Potentials of Solvated Molecules and Ions: Theory and Application to Aqueous Phenol and Phenolate计算溶剂化分子和离子的电离能和氧化还原电位的第一原理协议: 理论和对苯酚和酚盐水溶液的应用

