返回
Supervised pre-training for improved stability in deep reinforcement learning
DOI:10.1016/j.icte.2021.12.015.png)
摘要
En 中文
Deep reinforcement learning (DRL) technology has been actively studied with the recent advances in deep learning. As a result, the researchers are continuously improving performance and expanding the applications. However, recent literature reports that the performance of DRL is sensitive to the various design choices, e.g., the neural network initialization. Accordingly, it makes DRL hard to obtain a stable performance, which degrades reproducibility. Therefore, we propose a supervised pre-training method for both policy and value networks to improve stability. We pre-train the policy network to maximize the initial entropy and pre-train the value network to bias the distribution to a specific value. The experiments are conducted on tasks with discrete action space where it is hard to control the initial entropy. Through the experiments, the effectiveness of the proposed method in terms of stability and performance is validated.(c) 2021 The Author(s). Published by Elsevier B.V. on behalf of The Korean Institute of Communications and Information Sciences. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Keyword:
Deep reinforcement learning
Pre-training
Supervised learning
Maximum entropy
Stability
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
4.2
论文数:
1.0K
被引数:
2.5K
机构
引用论文
Study of bi-directional buck-boost converter topologies for application in electrical vehicle motor drives应用于电动汽车电机驱动的双向buck-boost变换器拓扑研究
Distributed deep reinforcement learning for autonomous aerial eVTOL mobility in drone taxi applications无人机出租车应用中自主空中eVTOL移动性的分布式深度强化学习
ICT EXPRESS
IF4.2
Deep Q-learning-based resource allocation for solar-powered users in cognitive radio networks
ICT EXPRESS
IF4.2
没有更多内容

