Return
MAD3PG: A Framework for Multi-Agent Deep Denoising Diffusion Policy Gradient Optimization
DOI:10.1016/j.inffus.2025.104026.png)
Abstract
En 中文
• Proposes MAD3PG, the first framework using denoising diffusion models for multi-agent value distribution estimation, and provides corresponding theoretical analysis. • Introduces a K-repeated sampling strategy with temporal-difference targets to enable efficient training of diffusion models in online reinforcement learning. • Demonstrates superior robustness and efficiency over MADDPG-type algorithms through extensive experiments in MPE and MuJoCo environments.
Journal
IF:
15.5
Papers:
4.1K
Citations:
2.7W

