arrow
返回

Two steps reinforcement learning

delete2008-01-01
delete20
delete
OA
AI
F
Fernando Fernández *
D
Daniel Borrajo
DOI:10.1002/int.20255delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
When applying reinforcement learning in domains with very large or continuous state spaces, the experience obtained by the learning agent in the interaction with the environment must be generalized. The generalization methods are usually based on the approximation of the value functions used to compute the action policy and tackled in two different ways. On the one hand by using an approximation of the value functions based on a supervized learning method. On the other hand, by discretizing the environment to use a tabular representation of the value functions. In this work, we propose an algorithm that uses both approaches to use the benefits of both mechanisms, allowing a higher performance. The approach is based on two learning phases. In the first one, a learner is used as a supervized function approximator, but using a machine learning technique which also outputs a state space discretization of the environment, such as nearest prototype classifiers or decision trees do. In the second learning phase, the space discretization computed in the first phase is used to obtain a tabular representation of the value function computed in the previous phase, allowing a tuning of such value function approximation. Experiments in different domains show that executing both learning phases improves the results obtained executing only the first one. The results take into account the resources used and the performance of the learned behavior. (c) 2008 Wiley Periodicals, Inc.
Keyword:
ALGORITHM
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

International Journal of Intelligent Systems 封面图
International Journal of Intelligent Systems
IF:
3.7
论文数:
3.0K
被引数:
8.1K

机构

U
Universidad Carlos III de Madrid
学者数:
5.5K
论文数: 5.7K
被引数: 4.5K
引用论文

引用论文

Towards energy-autonomous wake-up receiver using Visible Light Communication
err2016-01-01
err0
errOAAI
errJoyce Sariol Ramos; Ilker Demirkol; Josep Paradells; Daniel Vossing; Karim M. Gad; Martin Kasemann
err分享
err收藏
err1997-01-01
err0
PREAI
errLauren B. Alloy; Nancy Just; Catherine Panzarella
err分享
err收藏
Feature weighting in k-means clustering
err2003-01-01
err302
errOAAI
errModha, DS; Spangler, WS
err分享
err收藏
err分享
err收藏
err分享
err收藏
Microstructure of itaconate polymers containing benzyl groups
err2003-03-12
err0
PREAI
errArturo Horta; Irmina Herńndez‐Fuentes; Ligia Gargallo; Deodato Radić
err分享
err收藏
学者 查看更多内容