1
Return

Optimizing device-to-device path discovery in 5G networks using distributional dueling Q-learning with battery constraints

delete2026-03-01
delete0
PRE
AI
R
Rashmi *
K
Kumar, Prashant
DOI:10.1016/j.phycom.2026.103048delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Device-to-Device (D2D) communication is vital in Fifth-Generation (5G) networks for reducing latency and offloading base stations, but its effectiveness is constrained by two persistent challenges: finding optimal multi-hop routes in dynamic conditions and preserving device battery life. Existing routing schemes typically optimize one at the expense of the other, leading to inefficient paths or premature device shutdown. This paper introduces a Distributional Dueling Q-learning (2-DQ) algorithm that decomposes the action-value (Q) function into state value and action advantage terms while explicitly enforcing a 30% minimum battery threshold. Extensive simulations show that 2-DQ delivers a 23% gain in route efficiency, a 19% improvement in adaptability under dense and heterogeneous network scenarios, and a 17% boost in energy optimization compared to standard D2D and single-dueling Q-learning approaches. Moreover, the algorithm consistently maintains device battery levels above operational thresholds in urban, rural, and industrial testbeds. These results position 2-DQ as a scalable and energy-aware framework for real-time D2D path selection in next-generation 5G deployments.
Keywords:
D2D communication
Path optimization
Q-learning
Reinforcement learning
Battery life constraint
Reward maximization

Journal

Physical Communication cover
Physical Communication
IF:
2.2
Papers:
279
Citations:
2.6K

Organization

N
national institute of technology (nit system)
Scholars:
3.9W
Papers: 3.7W
Citations: 31
Cited Papers

Cited Papers

Citing Papers

Citing Papers