arrow
Return

Robust Offline Multi-Agent Reinforcement Learning for Latency-Aware SDN Path Control in 6G-Oriented Network Softwarization

delete2026-08-01
delete0
delete
OA
AI
А
Абзал Кызырканов *
Y
Yedil Nurakhov
Z
Zhenis Otarbay *
D
Danil Lebedev
DOI:10.3390/technologies14080468delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Future sixth-generation (6G)-oriented networks require programmable control that can adapt routing to latency and congestion without unsafe online exploration. This study evaluates offline multi-agent deep deterministic policy gradient (MADDPG) with behavior-adjusted training rewards for latency-aware path control in software-defined networking (SDN). Each traffic pair is modeled as an agent selecting one of three retained candidate paths, while centralized critics learn coordinated decisions from topology-specific Ryu–Mininet transition datasets. Nine policies are compared using ten paired seeds on fat-tree, mesh-grid, and WAN-corridors topologies under a deployed utilization–latency weighting of 0.60 / 0.40 , together with flow-completion, latency, congestion, architectural-comparison, sensitivity, robustness, statistical, and controller-overhead analyses. The utilization-aware path heuristic achieves the strongest overall reward ranking. MADDPG is the strongest learned policy on fat-tree, is not significantly outperformed by any evaluated policy on mesh-grid, and remains statistically tied with completion-matched policies on WAN-corridors. Behavior adjustment is topology-dependent rather than uniformly beneficial. The exported policy requires approximately 52 μ s per joint decision, whereas complete control-loop timing is dominated by network-statistics polling. These results support offline multi-agent SDN control as a competitive, low-overhead option when interpreted jointly with topology structure, flow completion, and strong heuristic baselines.
Keywords:
software-defined networking
offline reinforcement learning
multi-agent reinforcement learning
MADDPG
traffic engineering
path control
latency-aware routing
6G network softwarization

Journal

T
Technologies
IF:
3.6
Papers:
1.2K
Citations:
3.2K

Organization

A
Astana IT University
Scholars:
226
Papers: 123
Citations: 1
A
Al-Farabi Kazakh National University
Scholars:
2.9K
Papers: 1.5K
Citations: 1.6K