Return
A Decomposition Optimization-Based Multiobjective Reinforcement Learning Algorithm for Obtaining Nonconvex Pareto Fronts
DOI:10.1109/TNNLS.2025.3603165.png)
Abstract
En 中文
Multiobjective reinforcement learning (MORL) aims to seek a complete Pareto front (PF) with different compromise policies in multiobjective Markov decision processes (MOMDPs). However, most MORL algorithms currently have a limitation in handling the MOMDPs with nonconvex PFs. In this article, we propose a nonlinear MORL algorithm based on decomposition and variance reduction (MORL/D-VR) to overcome this limitation. MORL/D-VR adopts the Tchebycheff approach to transform a given MOMDP into a set of single-objective Markov decision processes (MDPs) and subsequently applies an improved policy gradient algorithm, called expected utility policy gradient (EUPG), to solve each single-objective MDP efficiently. We analyze the Pareto optimality of employing the Tchebycheff approach and policy gradient methods that use the full return to update policy for solving MOMDPs. The analysis shows that such a case can identify any Pareto optimal policy regardless of the shape of PFs theoretically. This can provide a theoretical guarantee for applying the Tchebycheff approach and EUPG in MORL/D-VR to obtain the policies within the nonconvex PFs. Moreover, we devise a new baseline for EUPG to reduce the variance of gradient updates and adopt a weight vector adaptation method to improve diversity. The experimental results show that MORL/D-VR achieves a desirable performance in handling problems with different convex and nonconvex PFs and outperforms current state-of-the-art MORL algorithms.
Keywords:
Multiobjective Markov decision process (MOMDP)
multiobjective optimization (MOO)
multiobjective reinforcement learning (MORL)
Pareto optimality
Tchebycheff approach
Journal
IF:
8.9
Papers:
7.5K
Citations:
7.2W

