Return
Delay-Optimal Random Access: A Learning Framework
DOI:10.1109/TCOMM.2026.3652495.png)
Abstract
En 中文
For supporting ever-growing demands on low-latency services in Machine-to-Machine (M2M) communications, it is crucial to optimize the delay performance of random access. To that end, much effort has been made on the optimal tuning of access parameters with a given access strategy. For further optimization of access strategy, learning-based approaches may have great potential, with which nodes learn the best access strategy based on their own observations and experience. Yet how to leverage learning-based approaches for delay optimization has remained largely unexplored. In this paper, three main challenges in delay-optimal learning-based access design are identified and tackled for sensing-free random access networks. First of all, the minimum mean queueing delay with central coordination is derived, which serves as a delay bound for the distributed case. For each node independently making decisions, a $Q$ -learning algorithm, DORA-GS, is proposed, and demonstrated to be able to achieve the delay bound if each node has the global information about the occupancy of all the others’ queues. When such information is unavailable, a Multi-Armed-Bandit (MAB)-based access scheme, DORA-MAB, is further developed, with which finite mean queueing delay can be achieved for the whole input rate region of $(0,1)$ . Gains over the existing representative sensing-free random access schemes are shown to be significant especially when the aggregate input rate is high.
Keywords:
Random access
queueing delay
Markov decision process
$Q$ -learning
multi-armed bandit
learning-based access
Journal
IF:
8.3
Papers:
1.2W
Citations:
3.6W

