Return
Provable policy optimization for attitude control systems with unknown inertia
J
F
Q
K
DOI:10.1016/j.automatica.2026.113233.png)
Abstract
En 中文
This paper proposes a provably convergent policy optimization (PO) method for rest-to-rest attitude reorientation of rigid bodies with unknown but bounded inertia. The control architecture integrates a PD-like feedback term with a nonlinear feedforward compensator. To prevent unwinding, a sign-dependent term is incorporated into the control design. Instead of requiring exact inertia values or conservative robustness margins, the gains of the structured controller, referred to as the policy, are directly learned from trajectory data by minimizing a quadratic performance cost. Moreover, a bounds-based initialization strategy is proposed to ensure that the initial policy lies within a locally convex and smooth region, thereby guaranteeing that the zeroth-order PO algorithm converges to the optimal policy with high probability. Simulation results verify the effectiveness and robustness of the proposed method.
Keywords:
Attitude control
Policy optimization
Unknown inertia
Anti-unwinding
Convergence guarantees
Journal
IF:
5.9
Papers:
1.1W
Citations:
5.2W
