arrow
返回

Decentralized graph-based multi-agent reinforcement learning using reward machines

delete2024-01-01
delete4
delete
OA
AI
J
Jueming Hu
Z
Zhe Xu
G
Guannan Qu
Y
Yutian Pang
Y
Yongming Liu *
DOI:10.1016/j.neucom.2023.126974delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
In multi-agent reinforcement learning (MARL), it is challenging for a collection of agents to learn complex temporally extended tasks. The difficulties lie in computational complexity and how to learn the high-level ideas behind reward functions. We study the graph-based Markov Decision Process (MDP), where the dynamics of neighboring agents are coupled. To learn complex temporally extended tasks, we use a reward machine (RM) to encode each agent's task and expose reward function internal structures. RM has the capacity to describe high-level knowledge and encode non-Markovian reward functions. We propose a decentralized learning algorithm to tackle computational complexity, called decentralized graph-based reinforcement learning using reward machines (DGRM), that equips each agent with a localized policy, allowing agents to make decisions independently based on the information available to the agents. DGRM uses the actor-critic structure, and we introduce the tabular Q-function for discrete state problems. We show that the dependency of the Q-function on other agents decreases exponentially as the distance between them increases. To further improve efficiency, we also propose the deep DGRM algorithm, using deep neural networks to approximate the Q-function and policy function to solve large-scale or continuous state problems. The effectiveness of the proposed DGRM algorithm is evaluated by three case studies, two wireless communication case studies with independent and dependent reward functions, respectively, and COVID-19 pandemic mitigation. Experimental results show that local information is sufficient for DGRM and agents can accomplish complex tasks with the help of RM. DGRM improves the global accumulated reward by 119% compared to the baseline in the case of COVID-19 pandemic mitigation.
Keyword:
Decentralized
Multi-agent
Reinforcement learning
Reward machine
Efficiency
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Neurocomputing 封面图
Neurocomputing
IF:
6.5
论文数:
2.5W
被引数:
6.5W

机构

A
Arizona State University
学者数:
2.7W
论文数: 2.5W
被引数: 4.2W
A
arizona state university-tempe
学者数:
1.5W
论文数: 1.2W
被引数: 13
引用论文

引用论文

err分享
err收藏
Priming
err2002-01-01
err0
PREAI
errAnthony D. Wagner; Wilma Koutstaal
err分享
err收藏
err分享
err收藏
学者 查看更多内容