返回
Voting-Based Multiagent Reinforcement Learning for Intelligent IoT
DOI:10.1109/JIOT.2020.3021017.png)
摘要
En 中文
The recent success of single-agent reinforcement learning (RL) in Internet of Things (IoT) systems motivates the study of multiagent RL (MARL), which is more challenging but more useful in large-scale IoT. In this article, we consider a voting-based MARL problem, in which the agents vote to make group decisions and the goal is to maximize the globally averaged returns. To this end, we formulate the MARL problem based on the linear programming form of the policy optimization problem and propose a primal-dual algorithm to obtain the optimal solution. We also propose a voting mechanism through which the distributed learning achieves the same sublinear convergence rate as centralized learning. In other words, the distributed decision making does not slow down the process of achieving global consensus on optimality. Finally, we verify the convergence of our proposed algorithm with numerical simulations and conduct case studies in practical multiagent IoT systems.
Keyword:
Convergence
Internet of Things
Optimization
Collaboration
Task analysis
Learning (artificial intelligence)
Games
Multiagent reinforcement learning (MARL)
primal– dual algorithm
voting mechanism
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
8.9
论文数:
1.4W
被引数:
7.8W
机构
引用论文
Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
MACHINE LEARNING
IF2.9
Multi-agent reinforcement learning as a rehearsal for decentralized planning多智能体强化学习作为分散计划的演练
NEUROCOMPUTING
IF6.5

