返回
Sample Complexity and Overparameterization Bounds for Temporal-Difference Learning With Neural Network Approximation
DOI:10.1109/TAC.2023.3234234.png)
摘要
En 中文
In this article, we study the dynamics of temporal-difference (TD) learning with neural network-based value function approximation over a general state space, namely, neural TD learning. We consider two practically used algorithms, projection-free and max-norm regularized neural TD learning, and establish the first convergence bounds for these algorithms. An interesting observation from our results is that max-norm regularization can dramatically improve the performance of TD learning algorithms in terms of sample complexity and overparameterization. The results in this work rely on a Lyapunov drift analysis of the network parameters as a stopped and controlled random process.
Keyword:
Neural networks
Approximation algorithms
Markov processes
Convergence
Complexity theory
Reinforcement learning
Kernel
reinforcement learning (RL)
stochastic approximation
temporal-difference (TD) learning
期刊
IF:
7
论文数:
1.3W
被引数:
6.7W
机构
引用论文
Comparison of purified psoralen-inactivated and formalin-inactivated dengue vaccines in mice and nonhuman primates
Vaccine
IF0
Frequent Activation of c-
kis
as a Transforming Gene in Fibrosarcomas Induced by Methylcholanthrene
Science
IF0

