返回
Multi-User Delay-Constrained Scheduling With Deep Recurrent Reinforcement Learning
DOI:10.1109/TNET.2024.3359911.png)
摘要
En 中文
Multi-user delay-constrained scheduling is a crucial challenge in various real-world applications, such as wireless communication, live streaming, and cloud computing. The scheduler must make real-time decisions to guarantee both delay and resource constraints simultaneously, without prior information on system dynamics that can be time-varying and challenging to estimate. Additionally, many practical scenarios suffer from partial observability issues due to sensing noise or hidden correlation. To address these challenges, we propose a deep reinforcement learning (DRL) algorithm called Recurrent Softmax Delayed Deep Double Deterministic Policy Gradient (RSD4) (https://github.com/hupihe/RSD4), which is a data-driven method based on a Partially Observed Markov Decision Process (POMDP) formulation. RSD4 guarantees resource and delay constraints by Lagrangian dual and delay-sensitive queues, respectively. It also efficiently handles partial observability with a memory mechanism enabled by the recurrent neural network (RNN). Moreover, it introduces user-level decomposition and node-level merging to support large-scale multihop scenarios. Extensive experiments on simulated and real-world datasets demonstrate that RSD4 is robust to system dynamics and partially observable environments and achieves superior performance over existing methods.
Keyword:
Dynamic scheduling
Optimal scheduling
System dynamics
Reinforcement learning
Optimization
Delays
Observability
Delay-constrained
scheduling
partial observability
deep reinforcement learning
期刊
I
IF:
3.6
论文数:
4.4K
被引数:
9.5K
机构
引用论文
Efficacy and safety of etoricoxib 30 mg and celecoxib 200 mg in the treatment of osteoarthritis in two identically designed, randomized, placebo-controlled, non-inferiority studies
Rheumatology
IF0
Gannet optimization algorithm : A new metaheuristic algorithm for solving engineering optimization problemsGannet优化算法: 一种新的解决工程优化问题的元启发式算法

