arrow
Return

Online attentive kernel-based temporal difference learning

delete2023-10-01
delete2
delete
OA
AI
X
Xingguo Chen *
G
Guang Yang
S
Shangdong Yang
H
Huihui Wang
S
Shaokang Dong
高扬 cover
高扬 (Yang Gao)
DOI:10.1016/j.knosys.2023.110902delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Kernel-based reinforcement learning has received increasing attention because it requires less prior knowledge linear approximation and neural networks. Online kernel-based updating, however, is hindered by the challenge of catastrophic forgetting or interference. Sparse representation is a key method to address this issue, but existing methods fail to satisfy four criteria: learnability, nonprior, nontruncation, and explicitness. In this paper, we present an attentive kernel-based value function approximation as a learnable, nonprior, nontruncated, and explicit sparse representation. We propose the online attentive kernel-based temporal difference (OAKTD) algorithm, which employs twotimescale optimization, and provide a convergence analysis for our proposed algorithm. Experimental results show that OAKTD outperforms online kernel-based TD learning algorithms, and the TD learning algorithm with Tile Coding on classical tasks, i.e., Mountain Car, Acrobot, CartPole and Puddle World. (c) 2023 Elsevier B.V. All rights reserved.
Keywords:
Online kernel-based reinforcement learning
Catastrophic forgetting or interference
Sparse representation
Attentive function
Two-timescale optimization
Stability analysis
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.2W
Citations:
4.5W

Organization

N
nanjing university
Scholars:
7.7W
Papers: 5.6W
Citations: 87