Return
Deep reinforcement learning for joint resource management in beyond-diagonal RIS-enhanced multi-cell THz-NOMA systems
M
M
Z
DOI:10.1016/j.compeleceng.2026.111346.png)
Abstract
En 中文
Terahertz (THz) communications offer promising bandwidth for fifth-generation and beyond communications, but suffer from severe path loss and blockages. To address this issue, non-orthogonal multiple access (NOMA) enhances spectral efficiency by serving multiple users simultaneously, while beyond-diagonal reconfigurable intelligent surfaces (BD-RIS) compensate for attenuation via dynamic inter-element couplings. However, jointly optimizing transmit power, user association, and BD-RIS phase shifts in multi-cell networks creates a highly non-convex and complex problem. To overcome this challenge, we propose resource-aware actor–critic (RAAC) framework, built upon the deep deterministic policy gradient algorithm. The RAAC framework tackles these complex constraints by employing a deterministic strategy for continuous power and phase-shift control, while handling discrete user association through relaxed embedding. By integrating constraint-aware penalty functions and off-policy learning, the RAAC technique efficiently explores the massive decision space. Simulation results show that the RAAC scheme achieves a maximum sum rate of 32 bps/Hz at 50 users, outperforming twin delayed deep deterministic policy gradient by 10% and soft actor–critic/proximal policy optimization by 22%, with a robust 2.8 bps/Hz gain over conventional diagonal RIS. The proposed RAAC framework provides a highly robust deep reinforcement learning platform for intelligent radio resource management.
Journal
C
IF:
4.9
Papers:
6.7K
Citations:
1.3W
