1
Return

Provable policy optimization for attitude control systems with unknown inertia

delete2026-08-10
delete0
PRE
AI
J
Junyu Yao
F
Feiran Zhao
Q
Qinglei Hu
K
Keyou You *
DOI:10.1016/j.automatica.2026.113233delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This paper proposes a provably convergent policy optimization (PO) method for rest-to-rest attitude reorientation of rigid bodies with unknown but bounded inertia. The control architecture integrates a PD-like feedback term with a nonlinear feedforward compensator. To prevent unwinding, a sign-dependent term is incorporated into the control design. Instead of requiring exact inertia values or conservative robustness margins, the gains of the structured controller, referred to as the policy, are directly learned from trajectory data by minimizing a quadratic performance cost. Moreover, a bounds-based initialization strategy is proposed to ensure that the initial policy lies within a locally convex and smooth region, thereby guaranteeing that the zeroth-order PO algorithm converges to the optimal policy with high probability. Simulation results verify the effectiveness and robustness of the proposed method.
Keywords:
Attitude control
Policy optimization
Unknown inertia
Anti-unwinding
Convergence guarantees

Journal

Automatica cover
Automatica
IF:
5.9
Papers:
1.1W
Citations:
5.2W

Organization

E
eth zürich
Scholars:
1.4K
Papers: 529
Citations: 1
B
beijing university of technology
Scholars:
4.3K
Papers: 1.5K
Citations: 0
B
Beihang University
Scholars:
5.0W
Papers: 4.0W
Citations: 37
T
tsinghua university
Scholars:
11.5W
Papers: 9.9W
Citations: 137
Cited Papers

Cited Papers

Citing Papers

Citing Papers