arrow
返回

Neural-network-based synchronous iteration learning method for multi-player zero-sum games

delete2017-06-01
delete43
PRE
AI
宋睿卓 封面图
宋睿卓 (Ruizhuo Song) *
Q
Qinglai Wei
B
Biao Song
DOI:10.1016/j.neucom.2017.02.051delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In this paper, a synchronous solution method for multi-player zero-sum games without system dynamics is established based on neural network. The policy iteration (PI) algorithm is presented to solve the Hamilton-Jacobi-Bellman (HJB) equation. It is proven that the obtained iterative cost function is convergent to the optimal game value. For avoiding system dynamics, off-policy learning method is given to obtain the iterative cost function, controls and disturbances based on Pl. Critic neural network (CNN), action neural networks (ANNs) and disturbance neural networks (DNNs) are used to approximate the cost function, controls and disturbances. The weights of neural networks compose the synchronous weight matrix, and the uniformly ultimately bounded (UUB) of the synchronous weight matrix is proven. Two examples are given to show that the effectiveness of the proposed synchronous solution method for multi-player ZS games. (C) 2017 Elsevier B.V. All rights reserved.
Keyword:
Adaptive dynamic programming
Approximate dynamic programming
Adaptive critic designs
Multi-player
Iteration learning
Neural network
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Neurocomputing 封面图
Neurocomputing
IF:
6.5
论文数:
2.5W
被引数:
6.5W

机构

C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
引用论文

引用论文

err分享
err收藏
Optimal model-free output synchronization of heterogeneous systems using off-policy reinforcement learning
err2016-09-01
err148
errOAAI
errModares, Hamidreza; Nageshrao, Subramanya P.; Lopes, Gabriel A. Delgado; Babuska, Robert; Lewis, Frank L.
err分享
err收藏
Globally optimal distributed cooperative control for general linear multi-agent systems
err2016-08-01
err29
PREAI
errFeng, Tao; Zhang, Huaguang; Luo, Yanhong; Liang, Hongjing
err分享
err收藏
Fuzzy-Based Goal Representation Adaptive Dynamic Programming基于模糊目标表示的自适应动态规划
err2016-10-01
err42
errOAAI
errTang, Yufei; He, Haibo; Ni, Zhen; Zhong, Xiangnan; Zhao, Dongbin; Xu, Xin
err分享
err收藏
学者 查看更多内容