arrow
返回

Off-policy based adaptive dynamic programming method for nonzero-sum games on discrete-time system

delete2020-08-01
delete6
PRE
AI
Y
Yinlei Wen
张化光 封面图
张化光 (Huaguang Zhang) *
H
He Ren
K
Kun Zhang
DOI:10.1016/j.jfranklin.2020.05.038delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In this paper, a novel model-free reinforcement learning method based on off-policy is introduced to solve nonzero-sum games of discrete-time linear systems. Compared with the traditional policy iteration (PI) method, which requires the knowledge of system dynamics, the proposed method can be trained by state data directly. Moreover, the traditional PI method is proved to be influenced by probing noises. In the analysis of the proposed method, the probing noises are specifically considered and proved to have no influence on the convergence. The solution of the optimal Nash equilibrium is deduced. It is also proved that the proposed algorithm can be applied in both online manner and offline manner. A simulation of the nonzero-sum games control problem on an F-16 aircraft discrete-time system is presented, and the results verify the effectiveness of the proposed algorithm. (c) 2020 The Franklin Institute. Published by Elsevier Ltd. All rights reserved.
Keyword:
H-INFINITY CONTROL
FAULT-TOLERANT CONTROL
NONLINEAR-SYSTEMS
LINEAR-SYSTEMS
CONTROLLER-DESIGN
LEARNING SOLUTION
ITERATION
ALGORITHM
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

J
Journal of the Franklin Institute-Engineering and Applied Mathematics
IF:
3.7
论文数:
6.4K
被引数:
1.5W

机构

N
northeastern university - china
学者数:
3.2W
论文数: 2.7W
被引数: 37
引用论文

引用论文

err
IF0
err
err0
PREAI
err
err分享
err收藏
Advancing social justice and racial equity in the public sector
err2018-07-18
err0
PREAI
errVanessa Lopez-Littleton; Brandi Blessett; Julie Burr
err分享
err收藏
A Robust Observer-Based Sensor Fault-Tolerant Control for PMSM in Electric Vehicles
err2016-12-01
err329
PREAI
errKommuri, Suneel Kumar; Defoort, Michael; Karimi, Hamid Reza; Veluvolu, Kalyana Chakravarthy
err分享
err收藏
学者 查看更多内容