arrow
Return

Multiagent Inductive Policy Optimization

delete2025-09-01
delete0
PRE
AI
H
Huang, Yubo
X
Xiaowei Zhao *
DOI:10.1109/TNNLS.2025.3601360delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Policy optimization methods are promising to tackle high-complexity reinforcement learning (RL) tasks with multiple agents. In this article, we derive a general trust region for policy optimization methods by considering the effect of subpolicy combinations among agents in multiagent environments. Based on this trust region, we propose an inductive objective to train the policy function, which can ensure agents learn monotonically improving policies. Furthermore, we observe that the policy always updates very weakly before falling into a local optimum. To address this, we introduce a cost regarding policy distance in the inductive objective to strengthen the motivation of agents to explore new policies. This approach strikes a balance during training, where the policy update step size remains within the constraints of the trust region, preventing excessive updates while avoiding getting stuck in local optima. Simulations on wind farm (WF) control tasks and two multiagent benchmarks demonstrate the high performance of the proposed multiagent inductive policy optimization (MAIPO) method.
Keywords:
Training
Space exploration
Optimization methods
Q-learning
Iterative methods
Wind farms
Probability distribution
Convergence
Learning systems
Decision making
Inductive optimization objective
multiagent reinforcement learning (RL)
trust region
wind farm (WF) control

Journal

IEEE Transactions on Neural Networks and Learning Systems cover
IEEE Transactions on Neural Networks and Learning Systems
IF:
8.9
Papers:
7.6K
Citations:
7.2W

Organization

U
University of Warwick
Scholars:
2.2W
Papers: 2.2W
Citations: 85
Cited Papers

Cited Papers

Counterfactual Multi-Agent Policy Gradients
err2018-04-29
err0
errOAAI
errJakob Foerster; Gregory Farquhar; Triantafyllos Afouras; Nantas Nardelli; Shimon Whiteson
errShare
errSave
Multi-agent deep reinforcement learning: a survey
err2021-04-15
err317
errOAAI
errGronauer, Sven; Diepold, Klaus
errShare
errSave
err
IF0
err
err0
errOAAI
err
errShare
errSave
Q-learning
err1992-05-01
err0
errOAAI
errChristopher J. C. H. Watkins; Peter Dayan
errShare
errSave
researcher View more