1
Return

A deep reinforcement learning method for container drayage transportation considering customer pairs

delete2026-05-12
delete0
delete
OA
AI
C
Chao Huang *
Y
Yinan Cui
X
Xiaoyang Zhou
B
Boyang Qu *
L
Li Yan
H
Hui Zhang
DOI:10.1093/jcde/qwag047delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Container drayage transportation serves as a critical link in global supply chains, yet truck capacity constraints and the complex interplay of multi-customer requirements often compromise drayage efficiency. These factors collectively increase fuel consumption and operational costs, posing significant challenges for logistics optimization. To address these issues, this article investigates a container drayage problem with customer pairs, where each pickup node corresponds to a delivery node. The optimization aims to minimize the trucks’ total fuel consumption. A mixed-integer nonlinear programming model is formulated on a graph-based representation to capture the coupling between task dependencies and truck states. To reduce computational complexity, we linearize the model by introducing several auxiliary variables. Recognizing the exponential growth of solution space in large-scale scenarios, we propose a deep reinforcement learning (DRL) method that integrates a Markov decision process, policy gradient optimization, and an attention mechanism. The method features a sequential decision-making system with an enhanced attention mechanism, a carefully designed cumulative reward function, and tailored training strategies. Specifically, the encoder efficiently extracts task features from depot, pickup, and delivery nodes, while the decoder optimizes feature fusion to guide task selection. Importantly, the model explicitly incorporates symmetry between customer pairs in both the encoder and decoder, thereby improving solution quality. Extensive experiments validate that the mathematical model, solved via Gurobi, obtains optimal solutions for small-scale instances within 1900 seconds, while the proposed DRL method achieves the same optimal solutions within 2700 seconds. For medium- and large-scale instances, DRL outperforms Gurobi, simulated annealing, and large neighborhood search, consistently delivering superior solutions within acceptable computation time, demonstrating strong generalization and robustness. Ablation studies further confirm the individual contributions of the encoder, decoder, and training strategy, with the full model achieving the best performance. These results underscore the potential of DRL as an effective tool for sustainable container drayage optimization.
Keywords:
Container drayage
Customer pairs
Fuel consumption minimization
Mixed-integer nonlinear programming
Deep reinforcement learning

Journal

Journal of Computational Design and Engineering cover
Journal of Computational Design and Engineering
IF:
6.1
Papers:
392
Citations:
3.2K

Organization

Z
zhongyuan university of technology
Scholars:
804
Papers: 241
Citations: 0
J
Jiangsu University of Science and Technology
Scholars:
5.6K
Papers: 2.0K
Citations: 263
Cited Papers

Cited Papers

Citing Papers

Citing Papers