返回
Iterative adaptive dynamic programming methods with neural network implementation for multi-player zero-sum games
DOI:10.1016/j.neucom.2018.04.005.png)
摘要
En 中文
This paper presents novel iterative learning methods along with the neural network implementation for multi-player zero-sum games. Solving zero-sum games depends on the solutions of Hamilton-Jacobi-Isaacs equations, which are nonlinear partial differential equations. These solutions are generally difficult or even impossible to be obtained analytically. To overcome this difficulty, iterative adaptive dynamic programming algorithms are utilized. In the related research works, three-network architecture, i.e., critic-actor-disturbance structure, is used to approximate the value function, control policies and disturbance policies. Different from the previous works, this paper employs single-network architecture, i.e., critic-only structure, to implement the proposed algorithms, which reduces the computation burden and the complexity of design procedure. Finally, two simulation examples are provided to illustrate the effectiveness of our proposed methods. (C) 2018 Published by Elsevier B.V.
Keyword:
Adaptive dynamic programming
Approximate dynamic programming
Zero-sum games
Neural networks
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6.5
论文数:
2.5W
被引数:
6.5W
机构
引用论文
Neural-network-based synchronous iteration learning method for multi-player zero-sum games基于神经网络的多人零和游戏同步迭代学习方法
NEUROCOMPUTING
IF6.5
Finite-Approximation-Error-Based Discrete-Time Iterative Adaptive Dynamic Programming基于有限近似误差的离散时间迭代自适应动态规划
H∞ Tracking Control of Completely Unknown Continuous-Time Systems via Off-Policy Reinforcement Learning基于非策略强化学习的完全未知连续时间系统的h ∞ 跟踪控制
Adaptive Dynamic Programming and Adaptive Optimal Output Regulation of Linear Systems线性系统的自适应动态规划和自适应最优输出调节
A three-network architecture for on-line learning and optimization based on adaptive dynamic programming
NEUROCOMPUTING
IF6.5

