arrow
返回

Efficient and stable deep reinforcement learning: selective priority timing entropy

delete2024-08-09
delete0
PRE
AI
L
Lin Huo
J
Jianlin Mao *
H
Hongjun San
S
Shufan Zhang
R
Ruiqi Li
L
Lixia Fu
DOI:10.1007/s10489-024-05705-6delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Deep reinforcement learning (DRL) has made significant strides in addressing tasks with high-dimensional continuous action spaces. However, the field still faces the challenges of low sample utilization and insufficient exploration-exploitation balance, limiting the generalizability of algorithms across different environments. To effectively improve sample utilization, optimize the exploration-exploitation balance, and achieve higher rewards in tasks, this paper designs the selective priority timing entropy (SPTE) algorithm. Subsequently, selective prioritized experience replay (SPER) is proposed, which employs frequent replay of multiframe memories to enhance sample utilization and improve the stability of policy updates. Additionally, the temporal advantage with decay (TAD) method introduces a decay factor to help adjust the weights of the variance and bias, thereby reducing estimation errors. The reward mechanism is augmented with multientropy (ME) for entropy-regularized training, achieving a balance between information exploration and exploitation. Finally, experimental testing on the challenging Arcade platform demonstrated that the SPTE algorithm surpasses the average testing level of human players by 104.936%. Furthermore, compared to other algorithms, SPTE achieves an average score increase of over 32.75%, and it consistently outperforms the compared methods in more than 60% of tasks, indicating its strong adaptability and robustness.
Keyword:
Deep reinforcement learning
Sample utilization
Exploration-exploitation
Selective prioritized experience replay
Temporal advantage with decay
Multientropy

期刊

Applied Intelligence 封面图
Applied Intelligence
IF:
3.5
论文数:
7.6K
被引数:
1.7W

机构

暂无机构信息
引用论文

引用论文

Evaluation of iron loading in four types of hepatopancreatic cells of the mangrove crab Ucides cordatus using ferrocene derivatives and iron supplements
err2018-03-27
err0
PREAI
errHector Aguilar Vitorino; Priscila Ortega; Roxana Y. Pastrana Alta; Flavia Pinheiro Zanotto; Breno Pannia Espósito
err分享
err收藏
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent Domain
err2024-07-01
err41
errOAAI
errHao, Jianye; Yang, Tianpei; Tang, Hongyao; Bai, Chenjia; Liu, Jinyi; Meng, Zhaopeng; Liu, Peng; Wang, Zhen
err分享
err收藏
err分享
err收藏
学者 查看更多内容