arrow
返回

Implementing action mask in proximal policy optimization (PPO) algorithm

delete2020-09-01
delete34
delete
OA
AI
C
Cheng-Yen Tang
C
Chien‐Hung Liu
W
Woei-Kae Chen
S
Shingchern D. You *
DOI:10.1016/j.icte.2020.05.003delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
The proximal policy optimization (PPO) algorithm is a promising algorithm in reinforcement learning. In this paper, we propose to add an action mask in the PPO algorithm. The mask indicates whether an action is valid or invalid for each state. Simulation results show that, when compared with the original version, the proposed algorithm yields much higher return with a moderate number of training steps. Therefore, it is useful and valuable to incorporate such a mask if applicable. (C) 2020 The Korean Institute of Communications and Information Sciences (KICS). Publishing services by Elsevier B.V.
Keyword:
PPO
Invalid action
Reinforcement learning
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

ICT Express 封面图
ICT Express
IF:
4.2
论文数:
1.0K
被引数:
2.5K

机构

N
National Taipei University of Technology
学者数:
7.1K
论文数: 7.3K
被引数: 6.8K
引用论文

引用论文