返回
Efficient Batch-Mode Reinforcement Learning Using Extreme Learning Machines
DOI:10.1109/TSMC.2019.2926806.png)
摘要
En 中文
As a class of batch-mode reinforcement learning (RL) methods for Markov decision problems with large or continuous state spaces, approximate policy iteration (API) has received increasing attention in the past decades. One open problem in the design of API algorithms is how to construct the basis functions or features for value function approximation (VFA). In this paper, we propose a novel batch-mode RL approach with randomly projected features for VFA. The proposed approach can be viewed as an extension of extreme learning machines (ELMs) to RL problems so it can be called ELM-API. The ELMs have been popularly studied in supervised learning problems, but there is not much work on the extension of ELMs to learning control problems. The proposed approach has advantages over the previous API algorithms in that the features for VFA can be quickly generated without complex parameter selection and the performance will be adaptive to different sample sets in batch-mode RL. In particular, the ELM-API approach can realize fast and efficient feature reconstruction when training sample sets are relatively small. Comprehensive simulation studies on two benchmark learning control problems were carried out to test the performance of API algorithms with different feature construction methods. It is shown that the ELM-API algorithm can obtain comparable or better performance than the previous API approaches. To further show the effectiveness of ELM-API in real-world applications, the simulation results on a more challenging high-dimensional lane-changing decision problem in dynamic traffic environment are also reported, which show the capability of the ELM-API algorithm in learning satisfactory lane-changing policies with high data efficiency.
Keyword:
Heuristic algorithms
Approximation algorithms
Kernel
Computer architecture
Prediction algorithms
Task analysis
Markov processes
Approximate policy iteration (API)
extreme learning machines (ELMs)
learning control
reinforcement learning (RL)
value function approximation (VFA)
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
10.5
论文数:
1.1W
被引数:
5.0W
机构
引用论文
On the performance of air-based solar heating systems utilizing phase-change energy storage
Energy
IF0
Productivity enhancement of solar still by PCM and Nanoparticles miscellaneous basin absorbing materials
Desalination
IF0
Study of bi-directional buck-boost converter topologies for application in electrical vehicle motor drives应用于电动汽车电机驱动的双向buck-boost变换器拓扑研究
A Clustering-Based Graph Laplacian Framework for Value Function Approximation in Reinforcement Learning基于聚类的图拉普拉斯框架,用于强化学习中的值函数逼近

