arrow
返回

Multi-User Delay-Constrained Scheduling With Deep Recurrent Reinforcement Learning

delete2024-06-01
delete2
PRE
AI
P
Pihe Hu
Y
Yu Chen
Z
Zhixuan Fang
F
Fu Xiao
L
Longbo Huang *
DOI:10.1109/TNET.2024.3359911delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Multi-user delay-constrained scheduling is a crucial challenge in various real-world applications, such as wireless communication, live streaming, and cloud computing. The scheduler must make real-time decisions to guarantee both delay and resource constraints simultaneously, without prior information on system dynamics that can be time-varying and challenging to estimate. Additionally, many practical scenarios suffer from partial observability issues due to sensing noise or hidden correlation. To address these challenges, we propose a deep reinforcement learning (DRL) algorithm called Recurrent Softmax Delayed Deep Double Deterministic Policy Gradient (RSD4) (https://github.com/hupihe/RSD4), which is a data-driven method based on a Partially Observed Markov Decision Process (POMDP) formulation. RSD4 guarantees resource and delay constraints by Lagrangian dual and delay-sensitive queues, respectively. It also efficiently handles partial observability with a memory mechanism enabled by the recurrent neural network (RNN). Moreover, it introduces user-level decomposition and node-level merging to support large-scale multihop scenarios. Extensive experiments on simulated and real-world datasets demonstrate that RSD4 is robust to system dynamics and partially observable environments and achieves superior performance over existing methods.
Keyword:
Dynamic scheduling
Optimal scheduling
System dynamics
Reinforcement learning
Optimization
Delays
Observability
Delay-constrained
scheduling
partial observability
deep reinforcement learning

期刊

I
IEEE-ACM Transactions on Networking
IF:
3.6
论文数:
4.4K
被引数:
9.5K

机构

T
tsinghua university
学者数:
11.9W
论文数: 10.0W
被引数: 137
引用论文

引用论文

A free-floating bike repositioning problem with faulty bikes
err2019-01-01
err0
errOAAI
errMuhammad Usama; Yongjun Shen; Onaira Zahoor
err分享
err收藏
Efficacy and safety of etoricoxib 30 mg and celecoxib 200 mg in the treatment of osteoarthritis in two identically designed, randomized, placebo-controlled, non-inferiority studies
err2007-01-25
err0
errOAAI
errC. O. Bingham; A. I. Sebba; B. R. Rubin; G. E. Ruoff; J. Kremer; S. Bird; S. S. Smugar; B. J. Fitzgerald; K. O'Brien; A. M. Tershakovec
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
Minimizing Latency for Data Aggregation inWireless Sensor Networks: An Algorithm Approach
err2022-08-30
err7
PREAI
errVan-Trung Pham; Nguyen, Tu N.; Liu, Bing-Hong; Thai, My T.; Dumba, Braulio; Lin, Tong
err分享
err收藏
A POMDP approach for scheduling the usage of airborne electronic countermeasures in air operations
err2016-01-01
err17
PREAI
errSong, Haifang; Xiao, Mingqing; Xiao, Jiyang; Liang, Yajun; Yang, Zhao
err分享
err收藏
Pulmonary Complications of Connective Tissue Diseases
err2008-03-01
err0
PREAI
errFelix Woodhead; Athol U. Wells; Sujal R. Desai
err分享
err收藏
学者 查看更多内容