Return
A Learning-Based Two-Stage Bidirectional Packing Framework for 3D Packing Problems
DOI:10.1109/TASE.2025.3607410.png)
Abstract
En 中文
The increasing demands of modern logistics have driven the need for the development of efficient packing methods capable of addressing the 3D Packing Problem (3D-PP). While Deep Reinforcement Learning (DRL) has emerged as a promising solution, the conventional three-stage scheme used in existing DRL-based methods still faces challenges, particularly in coordinating the behaviors of its constituent sub-networks and managing the large action space for item placement on instances involving large-sized bins. This work proposes a two-stage scheme to integrate item index and orientation selections into a single sub-stage, thereby simplifying behavioral coordination. To mitigate the issue of excessive memory usage associated with the selection integration, a Set Transformer with Induced Set Attention Block (ISAB) is employed to encode the rotated item state, thus keeping relatively light computation. Additionally, we propose a bidirectional packing method that compresses the placement action space while encouraging the agent to explore reasonable placement positions. Experimental results demonstrate that the Two-Stage Bidirectional Packing (TS-BP) framework, formed by the above components, improves space utilization of 2.8%~4.6% on high-difficulty packing instances compared to current state-of-the-art methods. The code is available at https://github.com/Ashenone511/Two-Stage-Bidirectional-Packing-Framework Note to Practitioners—This work proposes an effective learning-based Two-Stage Bidirectional Packing (TS-BP) framework to address the 3D Packing Problem (3D-PP), responding to the critical need to maximize warehouse storage density and reduce manufacturing logistics costs. The two-stage scheme equipped with the Induced Set Attention Block (ISAB) simplifies behavioral coordination among sub-networks while maintaining a lightweight computation. The bidirectional packing method reduces the action space size while ensuring the exploration of reasonable placement positions. The above components help to mitigate the shortcomings existing in current technologies. Simulation results demonstrate that the TS-BP framework outperforms state-of-the-art methods on standard and high-difficulty datasets. This work establishes a methodological framework for research on real-world packing solutions and contributes to improving warehouse storage density and reducing manufacturing logistics costs.
Keywords:
Deep reinforcement learning
3D packing problem
large action space
two-stage scheme
bidirectional packing
Journal
IF:
6.4
Papers:
4.9K
Citations:
1.6W

