arrow
Return

An Adaptive Exploration-Oriented Multi-Agent Co-Evolutionary Method Based on MATD3

delete2025-10-26
delete0
delete
OA
AI
S
Suyu Wang
Z
Zhentao Lyu
Q
Quan Yue
Q
Qichen Shang
Y
Ya Ke
F
Feng Gao *
DOI:10.3390/electronics14214181delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
As artificial intelligence continues to evolve, reinforcement learning (RL) has shown remarkable potential for solving complex sequential decision problems and is now applied in diverse areas, including robotics, autonomous vehicles, and financial analytics. Among the various RL paradigms, multi-agent reinforcement learning (MARL) stands out for its ability to manage cooperative and competitive interactions within multi-entity systems. However, mainstream MARL algorithms still face critical challenges in training stability and policy generalization due to factors such as environmental non-stationarity, policy coupling, and inefficient sample utilization. To mitigate these limitations, this study introduces an enhanced algorithm named MATD3_AHD, developed by extending the MATD3 framework, which integrates TD3 and MADDPG principles. The goal is to improve the learning efficiency and overall policy effectiveness of agents operating in complex environments. The proposed method incorporates three key mechanisms: (1) an Adaptive Exploration Policy (AEP), which dynamically adjusts the perturbation magnitude based on TD error to improve both exploration capability and training stability; (2) a Hierarchical Sampling Policy (HSP), which enhances experience utilization through sample clustering and prioritized replay; and (3) a Dynamic Delayed Update (DDU), which adaptively modulates the actor update frequency based on critic network errors, thereby accelerating convergence and improving policy stability. Experiments conducted on multiple benchmark tasks within the Multi-Agent Particle Environment (MPE) demonstrate the superior performance of MATD3_AHD compared to baseline methods such as MADDPG and MATD3. The proposed MATD3_AHD algorithm outperforms baseline methods—by an average of 5% over MATD3 and 20% over MADDPG—achieving faster convergence, higher rewards, and more stable policy learning, thereby confirming its robustness and generalization capability.
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

No journal information available

Organization