arrow
返回

PAC Reinforcement Learning Algorithm for General-Sum Markov Games

delete2023-05-01
delete2
delete
OA
AI
A
Ashkan Zehfroosh *
H
Herbert G. Tanner
DOI:10.1109/TAC.2022.3219340delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
This article presents a theoretical framework for probably approximately correct (PAC) multi-agent reinforcement learning (MARL) algorithms for Markov games. Using the idea of delayed Q-learning, this article extends the well-known Nash Q-learning algorithm to build a new PAC MARL algorithm for general-sum Markov games. In addition to guiding the design of a provably PAC MARL algorithm, the framework enables checking whether an arbitrary MARL algorithm is PAC. Comparative numerical results demonstrate the algorithm's performance and robustness.
Keyword:
Games
Markov processes
Picture archiving and communication systems
Nash equilibrium
Q-learning
Approximation algorithms
Convergence
Markov game
multiagent system
probably approximately correct (PAC)
reinforcement learning

期刊

IEEE Transactions on Automatic Control 封面图
IEEE Transactions on Automatic Control
IF:
7
论文数:
1.3W
被引数:
6.7W

机构

U
University of Delaware
学者数:
1.3W
论文数: 1.3W
被引数: 2.0W
引用论文

引用论文

err分享
err收藏
Temporal constraints on lens compensation in chicks
err2002-11-01
err0
errOAAI
errJonathan Winawer; Josh Wallman
err分享
err收藏
Antarctic climate change and the environment: an update南极气候变化与环境: 最新动态
err2013-04-18
err0
errOAAI
errJohn Turner; Nicholas E. Barrand; Thomas J. Bracegirdle; Peter Convey; Dominic A. Hodgson; Martin Jarvis; Adrian Jenkins; Gareth Marshall; Michael P. Meredith; Howard Roscoe; Jon Shanklin; John French; Hugues Goosse; Mauro Guglielmin; Julian Gutt; Stan Jacobs; Marlon C. Kennicutt; Valerie Masson-Delmotte; Paul Mayewski; Francisco Navarro; Sharon Robinson; Ted Scambos; Mike Sparrow; Colin Summerhayes; Kevin Speer; Alexander Klepikov
err分享
err收藏
err分享
err收藏
Growth and characterization of ceria thin films and Ce-doped γ-Al2O3nanowires using sol–gel techniques
err2010-10-26
err0
errOAAI
errS Gravani; K Polychronopoulou; V Stolojan; Q Cui; P N Gibson; S J Hinder; Z Gu; C C Doumanidis; M A Baker; C Rebholz
err分享
err收藏
学者 查看更多内容