arrow
返回

Sample Complexity and Overparameterization Bounds for Temporal-Difference Learning With Neural Network Approximation

delete2023-05-01
delete2
delete
OA
AI
S
Semih Çaycı *
S
Siddhartha Satpathi
N
Niao He
R
R. Srikant
DOI:10.1109/TAC.2023.3234234delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In this article, we study the dynamics of temporal-difference (TD) learning with neural network-based value function approximation over a general state space, namely, neural TD learning. We consider two practically used algorithms, projection-free and max-norm regularized neural TD learning, and establish the first convergence bounds for these algorithms. An interesting observation from our results is that max-norm regularization can dramatically improve the performance of TD learning algorithms in terms of sample complexity and overparameterization. The results in this work rely on a Lyapunov drift analysis of the network parameters as a stopped and controlled random process.
Keyword:
Neural networks
Approximation algorithms
Markov processes
Convergence
Complexity theory
Reinforcement learning
Kernel
reinforcement learning (RL)
stochastic approximation
temporal-difference (TD) learning

期刊

IEEE Transactions on Automatic Control 封面图
IEEE Transactions on Automatic Control
IF:
7
论文数:
1.3W
被引数:
6.7W

机构

R
RWTH Aachen University
学者数:
3.5W
论文数: 2.6W
被引数: 3.6W
U
University of Illinois Urbana-Champaign
学者数:
2.4W
论文数: 2.0W
被引数: 35
University of Illinois System 封面图
University of Illinois System
学者数:
6.8W
论文数: 6.2W
被引数: 644
学者 查看更多机构
引用论文

引用论文

Corticosterone impairs dendritic cell maturation and function
err2007-05-10
err0
errOAAI
errMichael D. Elftman; Christopher C. Norbury; Robert H. Bonneau; Mary E. Truckenmiller
err分享
err收藏
Comparison of purified psoralen-inactivated and formalin-inactivated dengue vaccines in mice and nonhuman primates
err2020-04-01
err0
errOAAI
errAppavu K. Sundaram; Daniel Ewing; Maria Blevins; Zhaodong Liang; Sandy Sink; Josef Lassan; Kanakatte Raviprakash; Gabriel Defang; Maya Williams; Kevin R. Porter; John W. Sanders
err分享
err收藏
学者 查看更多内容