返回
An efficient L2-norm regularized least-squares temporal difference learning algorithm
DOI:10.1016/j.knosys.2013.02.010.png)
摘要
En 中文
In reinforcement learning, when samples are limited in some real applications, Least-Squares Temporal Difference (LSTD) learning is prone to over-fitting, which can be overcome by the introduction of regularization. However, the solution of LSTD with regularization still depends on costly matrix inversion operations. In this paper we investigate the L2-norm regularized LSTD learning and propose an efficient algorithm to avoid expensive computational cost. We derive LSTD using Bellman operator along with projection operator. The L2-norm penalty is introduced to avoid over-fitting. We also describe the difference between Bellman residual minimization and LSTD. Then we propose an efficient recursive least-squares algorithm for L2-norm regularized LSTD, which can eliminate matrix inversion operations and decrease computational complexity effectively. We present empirical comparisons on the Boyan chain problem. The results show that the performance of the new algorithm is better than that of regularized LSTD. (C) 2013 Elsevier B.V. All rights reserved.
Keyword:
Reinforcement learning
Temporal difference
Recursive least-squares
Bellman residual minimizations
Regularization
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
K
IF:
7.6
论文数:
1.2W
被引数:
4.5W
机构
引用论文
An Analytical Model Based on Strain Localisation for the Study of Size‐Scale and Slenderness Effects in Uniaxial Compression Tests
Strain
IF0
Reinforcement learning of pedagogical policies in adaptive and intelligent educational systems自适应和智能教育系统中教学政策的强化学习
Echocardiographic assessment of abnormal left ventricular relaxation in man.超声心动图评估男性异常左心室舒张。
Heart
IF0


