arrow
Return

Reinforcement Learning-Based Policy Optimization for Heterogeneous Radio Access

delete2026-01-01
delete0
PRE
AI
A
Anup Mishra
Č
Čedomir Stefanović
P
Petar Popovski
I
Israel Leyva‐Mayorga
DOI:10.1109/LWC.2025.3642168delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Heterogeneous services such as broadband and latency-constrained Internet-of-Things (IoT), with diverse requirements, make efficient resource sharing challenging. Two canonical strategies are radio access network (RAN) Slicing, which allocates disjoint resources per service, and RAN Sharing, where services coexist on a common resource. Under either regime, grant-free repetition-based access is a key enabler of IoT connectivity, prompting latency-aware policies tuned to the regimes’ density and interference profiles. Prior work on repetition-policy optimisation mainly employs base station (BS)-centric reinforcement learning (RL) or irregular repetition slotted ALOHA (IRSA)-based approaches focused on throughput/packet loss rate (PLR) in asymptotic regimes. The former requires central orchestration, leading to high control overhead and reduced adaptability; the latter typically degrades when targeting performance at finite frame lengths. This motivates decentralised multi-agent RL (MARL), where each IoT device learns its own policy to meet latency targets in dense, finite frame-length, mixed-service deployments. To this end, we investigate an uplink scenario where a broadband user coexists with multiple latency-constrained IoT devices employing grant-free access. We formulate IoT access policy optimisation as a decentralised, model-free MARL problem, enabling devices to learn their transmission strategies under finite frame length structures and strict latency requirements. Under both RAN Slicing and RAN Sharing regimes, the proposed scheme outperforms baseline decentralised approaches; in each regime it achieves substantial IoT latency gains while sustaining broadband throughput and energy efficiency (EE). Furthermore, the results highlight and compare regime-specific trade-offs.
Keywords:
Heterogeneous 6G
Internet-of-Things (IoT)
reinforcement learning

Journal

I
IEEE Wireless Communications Letters
IF:
5.5
Papers:
665
Citations:
0

Organization

A
aalborg university
Scholars:
1.6W
Papers: 1.7W
Citations: 22