arrow
Return

Restricted gradient-descent algorithm for value-function approximation in reinforcement learning

delete2008-03-01
delete31
PRE
AI
A
André Barreto *
C
Charles W. Anderson
DOI:10.1016/j.artint.2007.08.001delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This work presents the restricted gradient-descent (RGD) algorithm, a training method for local radial-basis function networks specifically developed to be used in the context of reinforcement learning. The RGD algorithm can be seen as a way to extract relevant features from the state space to feed a linear model computing an approximation of the value function. Its basic idea is to restrict the way the standard gradient-descent algorithm changes the hidden units of the approximator, which results in conservative modifications that make the learning process less prone to divergence. The algorithm is also able to configure the topology of the network, an important characteristic in the context of reinforcement learning, where the changing policy may result in different requirements on the approximator structure. Computational experiments are presented showing that the RGD algorithm consistently generates better value-function approximations than the standard gradient-descent method, and that the latter is more susceptible to divergence. In the pole-balancing and Acrobot tasks, RGD combined with SARSA presents competitive results with other methods found in the literature, including evolutionary and recent reinforcement-learning algorithms. (c) 2007 Elsevier B.V. All rights reserved.
Keywords:
reinforcement learning
neuro-dynamic programming
value-function approximation
radial-basis-function networks

Journal

Artificial Intelligence Review cover
Artificial Intelligence Review
IF:
13.9
Papers:
6.1K
Citations:
1.9W

Organization

C
Colorado State University System
Scholars:
1.3W
Papers: 1.0W
Citations: 3
U
Universidade Federal do Rio de Janeiro
Scholars:
2.9W
Papers: 1.8W
Citations: 1.6W