arrow
Return

General multi-step value iteration for optimal learning control

delete2025-05-01
delete4
delete
OA
AI
王丁 (Ding Wang) *
W
Wang, Jiangyu
D
Derong Liu
J
Junfei Qiao
DOI:10.1016/j.automatica.2025.112168delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Learning control methods have been widely enhanced by reinforcement learning, but it is challenging to analyze the effects of incorporating extra system information. This paper presents a novel multi-step framework that utilizes extra multi-step system information to solve optimal control problems. Within this framework, we establish and classify general multi-step value iteration (MsVI) algorithms based on the uniformity between policy evaluation and improvement stages. According to this uniformity concept, the convergence condition and the acceleration conclusion are analyzed for different kinds of MsVI algorithms. Besides, we introduce a swarm policy optimizer to relieve limitations of the traditional gradient optimizer. Specifically, we implement general MsVI using an actor-critic scheme, where the swarm optimizer and neural networks are employed for policy improvement and evaluation, respectively. Furthermore, the approximation error caused by the approximator is also considered to verify the advantage of using multi-step system information. Finally, we apply the proposed method to a nonlinear benchmark system, demonstrating superior learning ability and control performance compared to traditional methods. (c) 2025 The Authors. Published by Elsevier Ltd. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Keywords:
Adaptive dynamic programming
Approximation error
Multi-step learning
Neural networks
Optimal control
Particle swarm optimization
Value iteration
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Automatica cover
Automatica
IF:
5.9
Papers:
1.1W
Citations:
5.2W

Organization

B
Beijing Univ Technol
Scholars:
2.6K
Papers: 1.2K
Citations: 354
S
southern univ sci technol
Scholars:
3.0K
Papers: 1.3K
Citations: 2