Return
Improving Energy Efficiency in Post-Disaster Networks: Multi-Agent Deep Reinforcement Learning for Enhancing Local Contributions
Y
H
T
X
Z
DOI:10.1109/tvt.2026.3666157.png)
Abstract
En 中文
To address the energy efficiency (EE) optimization challenge in post-disaster space-air enhanced integrated access and backhaul networks (SAE-IABN), a multi-agent global trust region local policy optimization (MA-GTRLPO) based resource allocation (RA) strategy is proposed. This strategy adopts a hybrid hierarchical architecture, decoupling global RA into centralized authorization and decentralized allocation to manage the multi-layer structured SAE-IABN’s resources efficiently. MA-GTRLPO enhances local contributions through a sequential local policy update rule, which helps stabilize convergence and avoids joint policy degradation risk. The general global performance difference bound theorem provides theoretical convergence guarantees for MA-GTRLPO. This theorem also offers a coordination mechanism for the Kullback-Leibler (KL) regularization coefficient and learning rate, improving global rewards by 6.9% compared to a heuristic hyperparameter baseline. Next, the global reward with the dynamic feedback framework and triple normalization is proposed to optimize EE further. The feedback framework with quality of service (QoS) constraints is introduced to prevent over-allocation. Then, the triple normalization helps agents clarify the contribution of EE, QoS guarantee, and energy constraint to the global reward. Simulations demonstrate that MA-GTRLPO outperforms the baselines in all metrics. It shows good scalability, robustness, and timeliness under varying network densities and topologies, guaranteeing QoS for user equipment (UE) and ensuring the endurance of uncrewed aerial vehicles (UAVs) remains above the minimum endurance bound.
Keywords:
Resource allocation
quality of service
multi-agent deep reinforcement learning
post-disaster communication
Journal
IF:
7.1
Papers:
1.7W
Citations:
6.6W
