Return
Sample efficient reinforcement learning via low-rank regularization
DOI:10.1016/j.knosys.2025.114176.png)
Abstract
En 中文
In this paper, the usefulness of low-rankness in state-action value function estimation is demonstrated using a simplified setup that is amenable to theoretical analysis. First, the concept of low-rank functions is defined motivated by standard functional analysis results. Subsequently, a specific procedure is proposed based on nuclear-norm penalized series estimation, in which the estimation of the low-rank function naturally leads to estimation of a low-rank matrix. Risk bounds are established for the estimator, which shows faster convergence rates compared to the standard estimator without using low-rankness. Several simulated toy examples are used as proof of concept to demonstrate the performances in simulations.
Keywords:
low-rankness
state-action value function
nuclear-norm penalized estimation
risk bounds
convergence rates
Journal
K
IF:
7.6
Papers:
1.3W
Citations:
4.5W
Organization
Cited Papers
No cited papers available

