arrow
返回

Voting-Based Multiagent Reinforcement Learning for Intelligent IoT

delete2021-02-15
delete10
delete
OA
AI
Y
Yue Xu
Z
Zengde Deng
M
Mengdi Wang
W
Wenjun Xu *
A
Anthony Man–Cho So
S
Shuguang Cui *
DOI:10.1109/JIOT.2020.3021017delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
The recent success of single-agent reinforcement learning (RL) in Internet of Things (IoT) systems motivates the study of multiagent RL (MARL), which is more challenging but more useful in large-scale IoT. In this article, we consider a voting-based MARL problem, in which the agents vote to make group decisions and the goal is to maximize the globally averaged returns. To this end, we formulate the MARL problem based on the linear programming form of the policy optimization problem and propose a primal-dual algorithm to obtain the optimal solution. We also propose a voting mechanism through which the distributed learning achieves the same sublinear convergence rate as centralized learning. In other words, the distributed decision making does not slow down the process of achieving global consensus on optimality. Finally, we verify the convergence of our proposed algorithm with numerical simulations and conduct case studies in practical multiagent IoT systems.
Keyword:
Convergence
Internet of Things
Optimization
Collaboration
Task analysis
Learning (artificial intelligence)
Games
Multiagent reinforcement learning (MARL)
primal– dual algorithm
voting mechanism
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Internet of Things Journal 封面图
IEEE Internet of Things Journal
IF:
8.9
论文数:
1.4W
被引数:
7.8W

机构

B
beijing university of posts & telecommunications
学者数:
1.4W
论文数: 1.2W
被引数: 9
T
The Chinese University of Hong Kong, Shenzhen
学者数:
4.3K
论文数: 4.0K
被引数: 7
S
Shenzhen Research Institute of Big Data
学者数:
256
论文数: 352
被引数: 357
M
ministry of education - china
学者数:
2.5W
论文数: 1.0W
被引数: 13
学者 查看更多机构
引用论文

引用论文

Distributed Policy Evaluation Under Multiple Behavior Strategies
err2015-05-01
err69
errOAAI
errValcarcel Macua, Sergio; Chen, Jianshu; Zazo, Santiago; Sayed, Ali H.
err分享
err收藏
Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
err2013-05-14
err158
errOAAI
errAzar, Mohammad Gheshlaghi; Munos, Remi; Kappen, Hilbert J.
err分享
err收藏
Distributed Reinforcement Learning via Gossip
err2017-03-01
err42
errOAAI
errMathkar, Adwaitvedant; Borkar, Vivek S.
err分享
err收藏
err分享
err收藏
学者 查看更多内容