返回
PAC Reinforcement Learning Algorithm for General-Sum Markov Games
DOI:10.1109/TAC.2022.3219340.png)
摘要
En 中文
This article presents a theoretical framework for probably approximately correct (PAC) multi-agent reinforcement learning (MARL) algorithms for Markov games. Using the idea of delayed Q-learning, this article extends the well-known Nash Q-learning algorithm to build a new PAC MARL algorithm for general-sum Markov games. In addition to guiding the design of a provably PAC MARL algorithm, the framework enables checking whether an arbitrary MARL algorithm is PAC. Comparative numerical results demonstrate the algorithm's performance and robustness.
Keyword:
Games
Markov processes
Picture archiving and communication systems
Nash equilibrium
Q-learning
Approximation algorithms
Convergence
Markov game
multiagent system
probably approximately correct (PAC)
reinforcement learning
期刊
IF:
7
论文数:
1.3W
被引数:
6.7W

