返回
Multiagent Inductive Policy Optimization
DOI:10.1109/TNNLS.2025.3601360.png)
摘要
En 中文
Policy optimization methods are promising to tackle high-complexity reinforcement learning (RL) tasks with multiple agents. In this article, we derive a general trust region for policy optimization methods by considering the effect of subpolicy combinations among agents in multiagent environments. Based on this trust region, we propose an inductive objective to train the policy function, which can ensure agents learn monotonically improving policies. Furthermore, we observe that the policy always updates very weakly before falling into a local optimum. To address this, we introduce a cost regarding policy distance in the inductive objective to strengthen the motivation of agents to explore new policies. This approach strikes a balance during training, where the policy update step size remains within the constraints of the trust region, preventing excessive updates while avoiding getting stuck in local optima. Simulations on wind farm (WF) control tasks and two multiagent benchmarks demonstrate the high performance of the proposed multiagent inductive policy optimization (MAIPO) method.
Keyword:
Training
Space exploration
Optimization methods
Q-learning
Iterative methods
Wind farms
Probability distribution
Convergence
Learning systems
Decision making
Inductive optimization objective
multiagent reinforcement learning (RL)
trust region
wind farm (WF) control
期刊
IF:
8.9
论文数:
7.6K
被引数:
7.2W
机构
引用论文
Demand-Responsive Transport Dynamic Scheduling Optimization Based on Multi-agent Reinforcement Learning Under Mixed Demand王, J.; 李, Y.; 孙, Q.; 汤, Y. 基于多智能体强化学习的混合需求下需求响应式交通动态调度优化. 在人工智能神经网络国际会议论文集, 卢加诺, 瑞士, 2024年9月17-20日; 施普林格自然: 朔姆, 瑞士, 2024. [谷歌学术] [CrossRef]
Integrated adaptive communication in multi-agent systems: Dynamic topology, frequency, and content optimization for efficient collaboration
NEUROCOMPUTING
IF6.5

