Return
Integrated Algorithm for PPO-Based Intelligent Dynamic Target Assignment and Entry Guidance Using Quasi-Second-Order Enhanced Linear Pseudospectral Method
DOI:10.1002/rnc.70190.png)
Abstract
En 中文
This paper presents an integrated algorithm for dynamic target assignment and entry guidance, which considers the dynamic variations in threat values of multiple targets and the strong nonlinear constraints of energy management during entry flight. This paper establishes an integrated computational framework by designing an energy management algorithm, an intelligent target assignment algorithm, and a guidance algorithm. It enables simultaneous target assignment and guidance command execution. Firstly, a high-precision reduced-order glide dynamics model, which more comprehensively accounts for the effects of Earth's rotation, is derived. This model, despite containing only three state variables, is capable of accurately predicting the glide trajectory of high-speed vehicles. To further reduce the solution scale, control is parameterized as longitudinal lift-to-drag ratio and energy at two reversal points. Subsequently, a quasi-second-order enhanced linear pseudospectral method is developed. This method initially expands the deviation propagation equation to the second order. To derive the analytical correction formula for control, the method further conducts Gaussian pseudospectral discretization and quasi-second-order linearization. Although the correction is more mathematically complex, it encompasses additional second-order information, which enables it to converge to higher precision within fewer iterations. Based on the energy management of the aircraft, the aforementioned work can be realized. To find a comprehensive optimal solution that meets the energy management requirements and maximizes the target value, the dynamic target assignment during entry is abstracted into a Markov decision process. The proximal policy optimization intelligent algorithm is employed for offline training of the assignment strategy, where the energy management algorithm is invoked to solve the energy management problem for each episode. The aircraft's energy and position, along with the threat values of each target, are utilized as observations, with the selected target number serving as the action. The reward function is designed as a composite function of target threat value, range-to-go, and energy management results. Finally, extensive numerical simulations are conducted to evaluate the algorithm. The results indicate that the aircraft can intelligently assign targets based on the remaining energy and the threat values of each target while strictly satisfying all constraints, which proves the algorithm's wide applicability, high precision, and fast computational efficiency. Furthermore, Monte Carlo simulation results indicate that even in highly dispersed environments, this integrated algorithm exhibits strong robustness.
Keywords:
dynamic target assignment
entry guidance
model predictive control
nonlinear control
pseudospectral method
reinforcement learning
Journal
IF:
3.2
Papers:
6.9K
Citations:
1.4W

