返回
Double distillation network for multi-agent reinforcement learning
DOI:10.1016/j.neucom.2026.133438.png)
摘要
En 中文
多智能体强化学习(MARL)通常在集中式训练和分布式执行(CTDE)框架下进行,以缓解环境非平稳性。然而,传统方法在执行过程中常受限于部分可观测性,导致智能体间出现累积间隙误差,阻碍了有效协同策略的发展。为解决此挑战,我们提出双蒸馏网络(DDN),其整合了两种互补的蒸馏模块,以在信息受限条件下促进鲁棒协调与高效协作。具体而言,外部蒸馏模块采用全局引导网络和局部策略网络,通过多级蒸馏促进知识转移,从而弥合全局训练与局部执行之间的差距。此外,内部蒸馏模块利用基于全局状态的内禀奖励增强智能体的探索能力。在StarCraft II和Predator–Prey基准上的实验结果表明,所提方法有效降低了固有误差,提升了探索效率,并显著改善了协同策略学习的质量。
Keyword:
Multi-agent reinforcement learning
Centralized training and decentralized execution
Knowledge distillation
Cooperative policy learning
Exploration efficiency
期刊
IF:
6.5
论文数:
2.5W
被引数:
6.5W
机构
引用论文
The Strategic Interplay Between the Platform's Store Brand Positioning and the Manufacturer's Core Category Innovation平台店铺品牌定位与制造商核心品类创新之间的战略互动
Mathematics
IF2.2
Symmetrical Learning and Transferring: Efficient Knowledge Distillation for Remote Sensing Image Classification
Symmetry
IF0
Multi-agent reinforcement learning as a rehearsal for decentralized planning多智能体强化学习作为分散计划的演练
NEUROCOMPUTING
IF6.5
Digital Twin Enhanced Federated Reinforcement Learning With Lightweight Knowledge Distillation in Mobile Networks移动网络中具有轻量级知识提炼的数字孪生增强联邦强化学习
Knowledge Distillation and Student-Teacher Learning for Visual Intelligence: A Review and New Outlooks视觉智能的知识蒸馏和师生学习: 回顾和新观点

