arrow
Return

Efficient exploration through active learning for value function approximation in reinforcement learning

delete2010-06-01
delete22
PRE
AI
T
Takayuki Akiyama *
H
Hirotaka Hachiya
M
Masashi Sugiyama
DOI:10.1016/j.neunet.2009.12.010delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Appropriately designing sampling policies is highly important for obtaining better control policies in reinforcement learning. In this paper, we first show that the least-squares policy iteration (LSPI) framework allows us to employ statistical active learning methods for linear regression. Then we propose a design method of good sampling policies for efficient exploration, which is particularly useful when the sampling cost of immediate rewards is high. The effectiveness of the proposed method, which we call active policy iteration (API), is demonstrated through simulations with a batting robot. (C) 2010 Elsevier Ltd. All rights reserved.
Keywords:
Reinforcement learning
Markov decision process
Least-squares policy iteration
Active learning
Batting robot
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Neural Networks cover
Neural Networks
IF:
6.3
Papers:
7.8K
Citations:
3.0W

Organization

I
Institute of Science Tokyo
Scholars:
3.2W
Papers: 2.7W
Citations: 117