返回
Accelerating deep reinforcement learning model for game strategy
DOI:10.1016/j.neucom.2019.06.110.png)
摘要
En 中文
In recent years, deep reinforcement learning has achieved impressing accuracies in games compared with traditional methods. Prior schemes utilized Convolutional Neural Networks (CNNs) or Long Short-Term Memory networks (LSTMs) to improve the performances of the agents. In this paper, we consider the issue from a different perspective when the training and inference of deep reinforcement learning are required to be performed with limited computing resources. Mainly, we propose two efficient neural network architectures of deep reinforcement learning: Light-Q-Network (LQN) and Binary-Q-Network (BQN). In LQN, The depth-wise separable CNNs are utilized in memory and computation saving. While, in BQN, the weights of convolutional layers are binary that help in shortening the training time and reduce memory consumption. We evaluate our approach on Atari 2600 domain and StarCraft II mini-games. The results demonstratethe efficiency of the proposed architectures. Though performances of agents in most games are still super-human, the proposed methods advance the agent from sub to super-human performance in particular games. Also, we empirically find that non-standard convolution and non-full-precision networks do not affect agent learning game strategy. (c) 2020 Elsevier B.V. All rights reserved.
Keyword:
Deep reinforcement learning
Convolutional neural network
Depthwise separable convolution
Binary weight network
期刊
IF:
6.5
论文数:
2.5W
被引数:
6.5W
机构
引用论文
Study of bi-directional buck-boost converter topologies for application in electrical vehicle motor drives应用于电动汽车电机驱动的双向buck-boost变换器拓扑研究
Fabrication of porous hollow γ-Al2O3 nanofibers by facile electrospinning and its application for water remediation静电纺丝法制备多孔中空 γ-Al2O3纳米纤维及其在水体修复中的应用
ICT for informal workers in Sub-Saharan Africa: Systematic review and analysis撒哈拉以南非洲非正规工人的信通技术: 系统回顾和分析
Adaptive Fuzzy Control of Strict-Feedback Nonlinear Time-Delay Systems with Full-State Constraints具有全状态约束的严格反馈非线性时滞系统的自适应模糊控制
没有更多内容

