返回
Neural-network-based synchronous iteration learning method for multi-player zero-sum games
DOI:10.1016/j.neucom.2017.02.051.png)
摘要
En 中文
In this paper, a synchronous solution method for multi-player zero-sum games without system dynamics is established based on neural network. The policy iteration (PI) algorithm is presented to solve the Hamilton-Jacobi-Bellman (HJB) equation. It is proven that the obtained iterative cost function is convergent to the optimal game value. For avoiding system dynamics, off-policy learning method is given to obtain the iterative cost function, controls and disturbances based on Pl. Critic neural network (CNN), action neural networks (ANNs) and disturbance neural networks (DNNs) are used to approximate the cost function, controls and disturbances. The weights of neural networks compose the synchronous weight matrix, and the uniformly ultimately bounded (UUB) of the synchronous weight matrix is proven. Two examples are given to show that the effectiveness of the proposed synchronous solution method for multi-player ZS games. (C) 2017 Elsevier B.V. All rights reserved.
Keyword:
Adaptive dynamic programming
Approximate dynamic programming
Adaptive critic designs
Multi-player
Iteration learning
Neural network
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6.5
论文数:
2.5W
被引数:
6.5W
机构
引用论文
Value iteration and adaptive dynamic programming for data-driven adaptive optimal control design数据驱动的自适应最优控制设计的值迭代和自适应动态规划
AUTOMATICA
IF5.9
Iterative GDHP-based approximate optimal tracking control for a class of discrete-time nonlinear systems
NEUROCOMPUTING
IF6.5
Optimal model-free output synchronization of heterogeneous systems using off-policy reinforcement learning
AUTOMATICA
IF5.9
Model-free multiobjective approximate dynamic programming for discrete-time nonlinear systems with general performance index functions
NEUROCOMPUTING
IF6.5
Globally optimal distributed cooperative control for general linear multi-agent systems
NEUROCOMPUTING
IF6.5
Finite-Approximation-Error-Based Discrete-Time Iterative Adaptive Dynamic Programming基于有限近似误差的离散时间迭代自适应动态规划
Adaptive optimal control for continuous-time linear systems based on policy iteration基于策略迭代的连续时间线性系统自适应最优控制
AUTOMATICA
IF5.9

