返回
Data-driven adaptive dynamic programming for continuous-time fully cooperative games with partially constrained inputs
DOI:10.1016/j.neucom.2017.01.076.png)
摘要
En 中文
In this paper, the fully cooperative game with partially constrained inputs in the continuous-time Markov decision process environment is investigated using a novel data-driven adaptive dynamic programming method. First, the model-based policy iteration algorithm with one iteration loop is proposed, where the knowledge of system dynamics is required. Then, it is proved that the iteration sequences of value functions and control policies can converge to the optimal ones. In order to relax the exact knowledge of the system dynamics, a model-free iterative equation is derived based on the model-based algorithm and the integral reinforcement learning. Furthermore, a data-driven adaptive dynamic programming is developed to solve the model-free equation using generated system data. From the theoretical analysis, we prove that this model-free iterative equation is equivalent to the model-based iterative equations, which means that the data-driven algorithm can approach the optimal value function and control policies. For the implementation purpose, three neural networks are constructed to approximate the solution of the model-free iteration equation using the off-policy learning scheme after the available system data is collected in the online measurement phase. Finally, two examples are provided to demonstrate the effectiveness of the proposed scheme. (C) 2017 Published by Elsevier B.V.
Keyword:
Adaptive dynamic programming
Optimal control
Neural network
Fully cooperative games
Data-driven
Constrained input
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6.5
论文数:
2.5W
被引数:
6.5W
机构
引用论文
Approximate N-Player Nonzero-Sum Game Solution for an Uncertain Continuous Nonlinear System不确定连续非线性系统的近似N人非零和博弈解
Integral reinforcement learning and experience replay for adaptive optimal control of partially-unknown constrained-input continuous-time systems
AUTOMATICA
IF5.9
Neural Network Based Online Simultaneous Policy Update Algorithm for Solving the HJI Equation in Nonlinear H∞ Control基于神经网络的在线同步策略更新算法,用于求解非线性h ∞ 控制中的HJI方程
Optimal control of unknown nonaffine nonlinear discrete-time systems based on adaptive dynamic programming基于自适应动态规划的未知非仿射非线性离散系统最优控制
AUTOMATICA
IF5.9

