arrow
Return

Decentralized Adaptive TD(λ) Learning With Linear Function Approximation: Nonasymptotic Analysis

delete2024-08-01
delete0
PRE
AI
J
Junlong Zhu
T
Tao Mao
M
Mingchuan Zhang *
葛泉波 cover
葛泉波 (Quanbo Ge)
Q
Qingtao Wu
李克勤 cover
李克勤 (Keqin Li)
DOI:10.1109/TSMC.2024.3382986delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In multiagent reinforcement learning, policy evaluation is a central problem. To solve this problem, decentralized temporal-difference (TD) learning is one of the most popular methods, which has been investigated in recent years. However, existing decentralized variants of TD learning often suffer from slow convergence due to the sensitive selection of learning rates. Inspired by the great success of adaptive gradient methods in the training of deep neural networks, this article proposes a decentralized adaptive TD (lambda) learning algorithm for general lambda with linear function approximation, referred to as D-AMSTD(lambda), which can mitigate the selective sensitivity of learning rates. Furthermore, we establish the finite-time performance bounds of D-AMSTD(lambda) under the Markovian observation model. The theoretical results show that D-AMSTD(lambda) can linearly converge to an arbitrarily small size of neighborhood of the optimal weight. Finally, we verify the efficacy of D-AMSTD(lambda) through a variety of experiments. The results show that D-AMSTD(lambda) outperforms existing decentralized TD learning methods.
Keywords:
Finite-time bounds
multiagent reinforcement learning (MARL)
policy evaluation
temporal-difference (TD) learning

Journal

IEEE Transactions on Cybernetics cover
IEEE Transactions on Cybernetics
IF:
10.5
Papers:
1.1W
Citations:
5.0W

Organization

S
state university of new york (suny) system
Scholars:
6.5W
Papers: 5.8W
Citations: 65