arrow
返回

Efficient Batch-Mode Reinforcement Learning Using Extreme Learning Machines

delete2021-06-01
delete8
PRE
AI
J
Jiahang Liu
左
左磊 (Lei Zuo)
徐鑫 封面图
徐鑫 (Xin Xu) *
X
Xinglong Zhang
J
Junkai Ren
Q
Qiang Fang
Xinwang Liu 封面图
Xinwang Liu (Xinwang Liu)
DOI:10.1109/TSMC.2019.2926806delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
As a class of batch-mode reinforcement learning (RL) methods for Markov decision problems with large or continuous state spaces, approximate policy iteration (API) has received increasing attention in the past decades. One open problem in the design of API algorithms is how to construct the basis functions or features for value function approximation (VFA). In this paper, we propose a novel batch-mode RL approach with randomly projected features for VFA. The proposed approach can be viewed as an extension of extreme learning machines (ELMs) to RL problems so it can be called ELM-API. The ELMs have been popularly studied in supervised learning problems, but there is not much work on the extension of ELMs to learning control problems. The proposed approach has advantages over the previous API algorithms in that the features for VFA can be quickly generated without complex parameter selection and the performance will be adaptive to different sample sets in batch-mode RL. In particular, the ELM-API approach can realize fast and efficient feature reconstruction when training sample sets are relatively small. Comprehensive simulation studies on two benchmark learning control problems were carried out to test the performance of API algorithms with different feature construction methods. It is shown that the ELM-API algorithm can obtain comparable or better performance than the previous API approaches. To further show the effectiveness of ELM-API in real-world applications, the simulation results on a more challenging high-dimensional lane-changing decision problem in dynamic traffic environment are also reported, which show the capability of the ELM-API algorithm in learning satisfactory lane-changing policies with high data efficiency.
Keyword:
Heuristic algorithms
Approximation algorithms
Kernel
Computer architecture
Prediction algorithms
Task analysis
Markov processes
Approximate policy iteration (API)
extreme learning machines (ELMs)
learning control
reinforcement learning (RL)
value function approximation (VFA)
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Cybernetics 封面图
IEEE Transactions on Cybernetics
IF:
10.5
论文数:
1.1W
被引数:
5.0W

机构

N
national university of defense technology - china
学者数:
1.8W
论文数: 1.4W
被引数: 9
引用论文

引用论文

Towards a wireless and fully-implantable ECoG system
err2013-06-01
err0
PREAI
errE. Tolstosheeva; J. Hoeffmann; J. Pistor; D. Rotermund; T. Schellenberg; D. Boll; T. Hertzberg; V. Gordillo-Gonzalez; S. Mandon; D. Peters-Drolshagen; M. Schneider; K. Pawelzik; A. Kreiter; S. Paul; W. Lang
err分享
err收藏
Opening Platforms: How, When and Why?
err2008-01-01
err0
PREAI
errThomas R. Eisenmann; Geoffrey Parker; Marshall W. Van Alstyne
err分享
err收藏
Hybrid least-squares algorithms for approximate policy evaluation
err2009-07-23
err18
errOAAI
errJohns, Jeff; Petrik, Marek; Mahadevan, Sridhar
err分享
err收藏
学者 查看更多内容