arrow
Return

Guided Proximal Policy Optimization with Structured Action Graph for Complex Decision-making

delete2025-01-07
delete0
PRE
AI
Y
Yiming Yang
邢登鹏 (Dengpeng Xing) *
W
Wannian Xia
王鹏 cover
王鹏 (Peng Wang)
DOI:10.1007/s11633-024-1503-7delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Reinforcement learning encounters formidable challenges when tasked with intricate decision-making scenarios, primarily due to the expansive parameterized action spaces and the vastness of the corresponding policy landscapes. To surmount these difficulties, we devise a practical structured action graph model augmented by guiding policies that integrate trust region constraints. Based on this, we propose guided proximal policy optimization with structured action graph (GPPO-SAG), which has demonstrated pronounced efficacy in refining policy learning and enhancing performance across sophisticated tasks characterized by parameterized action spaces. Rigorous empirical evaluations of our model have been performed on comprehensive gaming platforms, including the entire suite of StarCraft II and Hearthstone, yielding exceptionally favorable outcomes. Our source code is at https://github.com/sachiel321/GPPO-SAG.
Keywords:
Reinforcement learning
trust region policy optimization
complex decision-making
policy guiding
structured action graph

Journal

Machine Intelligence Research cover
Machine Intelligence Research
IF:
8.7
Papers:
301
Citations:
882

Organization

C
chinese academy of sciences
Scholars:
56.5W
Papers: 44.9W
Citations: 704