arrow
Return

Deep reinforcement learning algorithm incorporating problem characteristics for dynamic multi-objective permutation flow-shop scheduling problem

delete2025-07-01
delete0
PRE
AI
Y
Yuanyuan Yang
钱斌 cover
钱斌 (Bin Qian) *
胡蓉 cover
胡蓉 (Rong Hu)
李作成 cover
李作成 (Zuocheng Li)
金怀平 cover
金怀平 (Huaiping Jin)
J
Jian‐Bo Yang
DOI:10.1016/j.swevo.2025.101973delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Dynamic permutation flow shop scheduling problem (DPFSP) plays a critical role in real-world production systems, characterized by complex uncertainties including machine breakdowns, variable processing times, and the unpredictable job arrivals. Developing real-time solution approaches for such complex dynamic environments represents both a significant industrial need and a substantial computational challenge. Deep reinforcement learning (DRL) has demonstrated excellent capability for rapid and adaptive decision-making in complex dynamic environments, making it particularly suitable for DPFSP applications. This paper proposes a novel DRL algorithm incorporating problem characteristics (NDRLA_IPC) that specifically addresses the dynamic multiobjective PFSP (DMPFSP), with the objectives of minimizing the weighted maximum completion time and the total tardiness. NDRLA_IPC leverages a key insight that DMPFSP can be decomposed into a series of static PFSPs requiring real-time solutions, and implements a double deep Q-network (DDQN) architecture with components specifically engineered for DMPFSP characteristics. The algorithm introduces three key innovations: (1) a state feature vector design with high discrimination and generalization capabilities; (2) an action space designed to minimize temporal gaps between operations on the Gantt chart by leveraging processing constraints dynamically derived from the evolving problem state; and (3) a theoretically-validated reward function that effectively evaluates the online execution impact of each action. Comprehensive experiments demonstrate that NDRLA_IPC, after training on small-scale instances, transfers effectively to larger-scale DMPFSPs, delivering high-quality realtime solutions that outperform existing approaches across multiple performance metrics.
Keywords:
Deep reinforcement learning
Dynamic permutation flow shop scheduling
problem
Real-time scheduling
Multi-objective optimization

Journal

Swarm and Evolutionary Computation cover
Swarm and Evolutionary Computation
IF:
8.5
Papers:
2.1K
Citations:
1.0W

Organization

K
Kunming University of Science and Technology
Scholars:
9.1K
Papers: 2.5K
Citations: 2.1W
U
University of Manchester
Scholars:
5.7W
Papers: 5.2W
Citations: 7.4W