Return
Learning-Based Parallel Control for Unknown Nonaffine Nonzero-Sum Games
DOI:10.1109/TASE.2025.3593929.png)
Abstract
En 中文
In this paper, a novel nonzero-sum game (NSG) method is developed for completely unknown nonaffine nonlinear discrete-time (DT) systems, which is referred to as model-free NSG (MNSG). First, novel dynamic control laws are developed for NSGs using parallel control, namely introducing controls into feedback. Subsequently, an augmented N-player NSG is formulated according to the original N-player NSG to derive the dynamic control laws. Furthermore, we show that the control stabilities of the original and augmented N-player NSGs are equivalent. In the meantime, we prove that optimal control of the augmented N-player NSG is equivalent to near-optimal control of the original N-player NSG, and the Nash equilibrium of the original N-player NSG can be achieved. Then, a model-free learning scheme is developed to obtain the solution of the augmented N-player NSG using online policy iteration, and neither using a model network to predict unknown dynamics nor off-policy reinforcement learning (RL) is needed in the scheme. Lastly, numerical analysis, including the NSG of a DT system with unknown control-nonaffine dynamics and coupled controls, confirms the correctness of our MNSG method. The associated code is available at: https://github.com/lujingweihh/Adaptive-dynamic-programming-algorithms/tree/main/model_free_nonzero_sum_games_discrete_time Note to Practitioners—With the rise of information technologies and infrastructures, many real-world systems are now manipulated by multiple controllers, and the complexity of such systems is increasing. Hence, nonzero-sum games (NSGs) of multi-controller systems with general nonlinear dynamics, i.e., nonaffine dynamics, are an important research topic in academia and industry nowadays. Aiming at NSGs of multi-controller systems with unknown nonaffine dynamics, this paper develops a novel learning-based model-free NSG (MNSG) method using dynamic parallel control and adaptive dynamic programming (ADP). In the meantime, the developed learning-based MNSG method involves neither using a model network to predict unknown dynamics nor off-policy reinforcement learning (RL), contributing to a new and practical paradigm for MNSGs of multi-controller systems.
Keywords:
Adaptive dynamic programming
model-free optimal control
nonaffine dynamics
nonzero-sum games
parallel control
Journal
IF:
6.4
Papers:
4.9K
Citations:
1.6W

