Return
Off-policy based adaptive dynamic programming method for nonzero-sum games on discrete-time system
DOI:10.1016/j.jfranklin.2020.05.038.png)
Abstract
En 中文
In this paper, a novel model-free reinforcement learning method based on off-policy is introduced to solve nonzero-sum games of discrete-time linear systems. Compared with the traditional policy iteration (PI) method, which requires the knowledge of system dynamics, the proposed method can be trained by state data directly. Moreover, the traditional PI method is proved to be influenced by probing noises. In the analysis of the proposed method, the probing noises are specifically considered and proved to have no influence on the convergence. The solution of the optimal Nash equilibrium is deduced. It is also proved that the proposed algorithm can be applied in both online manner and offline manner. A simulation of the nonzero-sum games control problem on an F-16 aircraft discrete-time system is presented, and the results verify the effectiveness of the proposed algorithm. (c) 2020 The Franklin Institute. Published by Elsevier Ltd. All rights reserved.
Keywords:
H-INFINITY CONTROL
FAULT-TOLERANT CONTROL
NONLINEAR-SYSTEMS
LINEAR-SYSTEMS
CONTROLLER-DESIGN
LEARNING SOLUTION
ITERATION
ALGORITHM
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
J
IF:
3.7
Papers:
6.4K
Citations:
1.5W
Organization
Cited Papers
Iterative adaptive dynamic programming methods with neural network implementation for multi-player zero-sum games
NEUROCOMPUTING
IF6.5
Multi-player non-zero-sum games: Online adaptive learning solution of coupled Hamilton-Jacobi equations
AUTOMATICA
IF5.9

