Return
Benign Nonconvex Landscapes in Optimal and Robust Control, Part I: Global Optimality
Y
C
Y
DOI:10.1109/tac.2026.3675544.png)
Abstract
En 中文
Direct policy search has achieved great empirical success in reinforcement learning. Many recent studies have revisited its theoretical foundation for continuous control, which reveals elegant nonconvex geometry in various benchmark problems. This article considers two fundamental optimal and robust control problems with <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">partial observability:</i> linear quadratic Gaussian (LQG) control and <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$\mathcal {H}_\infty$</tex-math></inline-formula> robust control. In the policy space, the former problem is smooth but nonconvex, while the latter one is nonsmooth and nonconvex. We highlight some interesting and surprising “discontinuity” of LQG and <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$\mathcal {H}_\infty$</tex-math></inline-formula> cost functions around the boundary of their domains. Despite the lack of convexity (and possibly smoothness), we show that for a class of <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">nondegenerate</i> policies, all Clarke stationary points are globally optimal and there is no spurious local minimum for both LQG and <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$\mathcal {H}_\infty$</tex-math></inline-formula> control. The main results are established by a new and unified framework of Extended Convex Lifting (ECL), which reconciles the gap between nonconvex policy optimization and convex reformulations. This ECL framework is of independent interest, and we discuss its details in Part II of this article.
Keywords:
Benign nonconvexity
global optimality
linear quadratic Gaussian (LQG) control
nonconvex optimization
optimal and robust control
policy optimization
Journal
IF:
7
Papers:
1.3W
Citations:
6.7W
