Return
Dual-Critic Proximal Policy Optimization for Trajectory Tracking of Autonomous Underwater Vehicle Under Time-Varying Dynamics
Y
D
张
DOI:10.1109/joe.2026.3692007.png)
Abstract
En 中文
The autonomous underwater vehicle (AUV) serves as an ideal platform for replacing human operators in high-risk underwater operations, where high-precision trajectory tracking control is crucial for ensuring mission effectiveness. However, the time-varying dynamics caused by the complexity of the underwater environment present significant challenges to achieving high-precision trajectory tracking for AUVs. To address this challenge, a model-assisted deep reinforcement learning-based trajectory tracking method for an AUV with time-varying dynamics is presented in this article. Specifically, a novel state space is proposed, which replaces the absolute position of the agent with the relative position. This representation provides a more comprehensive description of the relationship between the AUV and the reference trajectory. The policy trained based on this state space can be easily transferred and applied to a different type of reference trajectory. Furthermore, by decomposing the reward function into three components, the sparse reward issue is effectively mitigated. The smoothness term within the reward function effectively suppresses abrupt changes in control commands, ensuring that the AUV maintains smooth navigation even under time-varying dynamics. Finally, an improved proximal policy optimization algorithm is proposed. The algorithm includes two critics that separately evaluate the value functions of actions under time-varying and static dynamics, and adaptively adjust their weights based on the intensity of the time-varying dynamics, effectively enhancing the policy’s disturbance rejection capability and robustness. Experimental results demonstrate that the proposed method maintains high-precision trajectory tracking across various time-varying dynamics, significantly outperforming other algorithms and validating its effectiveness.
Keywords:
Autonomous underwater vehicle (AUV)
deep reinforcement learning (DRL)
proximal policy optimization (PPO)
time-varying dynamics
trajectory tracking
Journal
IF:
5.3
Papers:
2.6K
Citations:
7.4K
