arrow
返回

Supervised pre-training for improved stability in deep reinforcement learning

delete2023-02-01
delete3
delete
OA
AI
S
Sooyoung Jang
H
Hyung-Il Kim *
DOI:10.1016/j.icte.2021.12.015delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Deep reinforcement learning (DRL) technology has been actively studied with the recent advances in deep learning. As a result, the researchers are continuously improving performance and expanding the applications. However, recent literature reports that the performance of DRL is sensitive to the various design choices, e.g., the neural network initialization. Accordingly, it makes DRL hard to obtain a stable performance, which degrades reproducibility. Therefore, we propose a supervised pre-training method for both policy and value networks to improve stability. We pre-train the policy network to maximize the initial entropy and pre-train the value network to bias the distribution to a specific value. The experiments are conducted on tasks with discrete action space where it is hard to control the initial entropy. Through the experiments, the effectiveness of the proposed method in terms of stability and performance is validated.(c) 2021 The Author(s). Published by Elsevier B.V. on behalf of The Korean Institute of Communications and Information Sciences. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Keyword:
Deep reinforcement learning
Pre-training
Supervised learning
Maximum entropy
Stability
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

ICT Express 封面图
ICT Express
IF:
4.2
论文数:
1.0K
被引数:
2.5K

机构

引用论文

引用论文

Molecular Subsets in the Gene Expression Signatures of Scleroderma Skin
err2008-07-16
err0
errOAAI
errAusra Milano; Sarah A. Pendergrass; Jennifer L. Sargent; Lacy K. George; Timothy H. McCalmont; M. Kari Connolly; Michael L. Whitfield
err分享
err收藏
Implementing action mask in proximal policy optimization (PPO) algorithm
err2020-09-01
err34
errOAAI
errTang, Cheng-Yen; Liu, Chien-Hung; Chen, Woei-Kae; You, Shingchern D.
err分享
err收藏
Deep Q-learning-based resource allocation for solar-powered users in cognitive radio networks
err2021-03-01
err9
errOAAI
errGiang, Hoang Thi Huong; Thanh, Pham Duy; Koo, Insoo
err分享
err收藏
Brain tumours in two Bactrian camels: a histiocytic sarcoma and a meningioma双峰驼脑肿瘤两例:一例组织细胞肉瘤和一例脑膜瘤
err2009-05-30
err0
errOAAI
errF. M. Molenaar; A. C. Breed; E. J. Flach; I. A. P. McCandlish; A. M. Pocknell; T. Strike; A. Routh; M. Taema; B. A. Summers
err分享
err收藏
没有更多内容