返回
Discrete-Time Stable Generalized Self-Learning Optimal Control With Approximation Errors
DOI:10.1109/TNNLS.2017.2661865.png)
摘要
En 中文
In this paper, a generalized policy iteration (GPI) algorithm with approximation errors is developed for solving infinite horizon optimal control problems for nonlinear systems. The developed stable GPI algorithm provides a general structure of discrete-time iterative adaptive dynamic programming algorithms, by which most of the discrete-time reinforcement learning algorithms can be described using the GPI structure. It is for the first time that approximation errors are explicitly considered in the GPI algorithm. The properties of the stable GPI algorithm with approximation errors are analyzed. The admissibility of the approximate iterative control law can be guaranteed if the approximation errors satisfy the admissibility criteria. The convergence of the developed algorithm is established, which shows that the iterative value function is convergent to a finite neighborhood of the optimal performance index function, if the approximate errors satisfy the convergence criterion. Finally, numerical examples and comparisons are presented.
Keyword:
Adaptive critic designs
adaptive dynamic programming (ADP)
approximate dynamic programming
generalized policy iteration (GPI)
neural networks
neurodynamic programming
nonlinear systems
optimal control
reinforcement learning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
8.9
论文数:
7.6K
被引数:
7.2W
机构
引用论文
Off-Policy Actor-Critic Structure for Optimal Control of Unknown Systems With Disturbances具有扰动的未知系统最优控制的非策略参与者-批评者结构
Infinite Horizon Self-Learning Optimal Control of Nonaffine Discrete-Time Nonlinear Systems非仿射离散非线性系统的无穷水平自学习最优控制
Discrete-Time Local Value Iteration Adaptive Dynamic Programming: Convergence Analysis离散局部值迭代自适应动态规划: 收敛性分析
Finite-Approximation-Error-Based Discrete-Time Iterative Adaptive Dynamic Programming基于有限近似误差的离散时间迭代自适应动态规划
Reinforcement-Learning-Based Robust Controller Design for Continuous-Time Uncertain Nonlinear Systems Subject to Input Constraints基于强化学习的具有输入约束的连续时间不确定非线性系统的鲁棒控制器设计
H∞ Tracking Control of Completely Unknown Continuous-Time Systems via Off-Policy Reinforcement Learning基于非策略强化学习的完全未知连续时间系统的h ∞ 跟踪控制

