arrow
Return

Provable distributed adaptive temporal-difference learning over time-varying networks*

delete2023-10-01
delete2
PRE
AI
J
Junlong Zhu
B
B. Li
L
Lin Wang *
M
Mingchuan Zhang
L
Ling Xing
J
Jiangtao Xi
Q
Qingtao Wu
DOI:10.1016/j.eswa.2023.120406delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Multi-agent reinforcement learning (MARL) has been successfully applied in many fields. In MARL, the policy evaluation problem is one of crucial problems. In order to solve this problem, distributed Temporal-Difference (TD) learning algorithm is one of the most popular methods in a cooperative manner. Despite its empirical success, however, the theory of the adaptive variant of distributed TD learning still remain limited. To fill this gap, we propose an adaptive distributed temporal-difference algorithm (referred to as MS-ADTD) under Markovian sampling over time-varying networks. Furthermore, we rigorously analyze the convergence of MS-ADTD, the theoretical results show that the local estimation can converge linearly to the optimal neighborhood. Meanwhile, the theoretical results are verified by simulation experiments.
Keywords:
Adaptive algorithms
Distributed temporal-difference
Markovian sampling
Non-asymptotic performance

Journal

Expert Systems with Applications cover
Expert Systems with Applications
IF:
7.5
Papers:
2.9W
Citations:
10.2W

Organization

U
University of Wollongong
Scholars:
1.3W
Papers: 1.6W
Citations: 2.8W