arrow
返回

Revisiting Approximate Dynamic Programming and its Convergence

delete2014-12-01
delete108
PRE
AI
A
Ali Heydari *
DOI:10.1109/TCYB.2014.2314612delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Value iteration-based approximate/adaptive dynamic programming (ADP) as an approximate solution to infinite-horizon optimal control problems with deterministic dynamics and continuous state and action spaces is investigated. The learning iterations are decomposed into an outer loop and an inner loop. A relatively simple proof for the convergence of the outer-loop iterations to the optimal solution is provided using a novel idea with some new features. It presents an analogy between the value function during the iterations and the value function of a fixed-final-time optimal control problem. The inner loop is utilized to avoid the need for solving a set of nonlinear equations or a nonlinear optimization problem numerically, at each iteration of ADP for the policy update. Sufficient conditions for the uniqueness of the solution to the policy update equation and for the convergence of the inner-loop iterations to the solution are obtained. Afterwards, the results are formed as a learning algorithm for training a neurocontroller or creating a look-up table to be used for optimal control of nonlinear systems with different initial conditions. Finally, some of the features of the investigated method are numerically analyzed.
Keyword:
Approximate dynamic programming
nonlinear control systems
optimal control
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Cybernetics 封面图
IEEE Transactions on Cybernetics
IF:
10.5
论文数:
1.1W
被引数:
5.0W

机构

暂无机构信息
引用论文

引用论文

err分享
err收藏
Do children have a theory of race?
err1995-02-01
err0
PREAI
errLawrence A. Hirschfeld
err分享
err收藏
学者 查看更多内容