Return
Optimizing device-to-device path discovery in 5G networks using distributional dueling Q-learning with battery constraints
R
K
DOI:10.1016/j.phycom.2026.103048.png)
Abstract
En 中文
Device-to-Device (D2D) communication is vital in Fifth-Generation (5G) networks for reducing latency and offloading base stations, but its effectiveness is constrained by two persistent challenges: finding optimal multi-hop routes in dynamic conditions and preserving device battery life. Existing routing schemes typically optimize one at the expense of the other, leading to inefficient paths or premature device shutdown. This paper introduces a Distributional Dueling Q-learning (2-DQ) algorithm that decomposes the action-value (Q) function into state value and action advantage terms while explicitly enforcing a 30% minimum battery threshold. Extensive simulations show that 2-DQ delivers a 23% gain in route efficiency, a 19% improvement in adaptability under dense and heterogeneous network scenarios, and a 17% boost in energy optimization compared to standard D2D and single-dueling Q-learning approaches. Moreover, the algorithm consistently maintains device battery levels above operational thresholds in urban, rural, and industrial testbeds. These results position 2-DQ as a scalable and energy-aware framework for real-time D2D path selection in next-generation 5G deployments.
Keywords:
D2D communication
Path optimization
Q-learning
Reinforcement learning
Battery life constraint
Reward maximization
Journal
IF:
2.2
Papers:
279
Citations:
2.6K
