arrow
Return

Distributed Off-Policy Temporal Difference Learning Using Primal-Dual Method

delete2022-01-01
delete0
delete
OA
AI
D
Donghwan Lee *
D
Do Wan Kim
J
Jianghai Hu
DOI:10.1109/ACCESS.2022.3211395delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The goal of this paper is to provide theoretical analysis and additional insights on a distributed temporal-difference (TD)-learning algorithm for the multi-agent Markov decision processes (MDPs) via saddle-point viewpoints. The (single-agent) TD-learning is a reinforcement learning (RL) algorithm for evaluating a given policy based on reward feedbacks. In multi-agent settings, multiple RL agents concurrently behave, and each agent receives its local rewards. The goal of each agent is to evaluate a given policy corresponding to the global reward, which is an average of the local rewards by sharing learning parameters through random network communications. In this paper, we propose a distributed TD-learning based on saddle-point frameworks, and provide rigorous analysis of finite-time convergence of the algorithm and its solution based on tools in optimization theory. The results in this paper provide general and unified perspectives of the distributed policy evaluation problem, and theoretically complement the previous works.
Keywords:
Convergence
Linear programming
Optimization
Markov processes
Symmetric matrices
Communication networks
Reinforcement learning
Machine learning
Sequential analysis
Multi-agent systems
Optimal control
Distributed processing
Reinforcement learning (RL)
multi-agent systems
convergence
temporal difference (TD) learning
machine learning
primal-dual method

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

Hanbat National University cover
Hanbat National University
Scholars:
2.0K
Papers: 2.2K
Citations: 1.9K
Purdue University System cover
Purdue University System
Scholars:
3.9W
Papers: 3.6W
Citations: 66