返回
Kernel-Based Decentralized Policy Evaluation for Reinforcement Learning
DOI:10.1109/TNNLS.2024.3453036.png)
摘要
En 中文
We investigate the decentralized nonparametric policy evaluation problem within reinforcement learning (RL), focusing on scenarios where multiple agents collaborate to learn the state-value function using sampled state transitions and privately observed rewards. Our approach centers on a regression-based multistage iteration technique employing infinite-dimensional gradient descent (GD) within a reproducing kernel Hilbert space (RKHS). To make computation and communication more feasible, we employ Nystrom approximation to project this space into a finite-dimensional one. We establish statistical error bounds to describe the convergence of value function estimation, marking the first instance of such analysis within a fully decentralized nonparametric framework. We compare the regression-based method to the kernel temporal difference (TD) method in some numerical studies.
Keyword:
Estimation
Kernel
Convergence
Hilbert space
Function approximation
Approximation algorithms
Vectors
Gradient descent (GD)
multiagent reinforcement learning (MARL)
policy iteration
reinforcement learning (RL)
reproducing kernel Hilbert space (RKHS)
state-value function
期刊
IF:
8.9
论文数:
7.6K
被引数:
7.2W
机构
引用论文
暂无论文信息

