arrow
Return

An adaptive multi-objective multi-task scheduling method by hierarchical deep reinforcement learning

delete2024-03-01
delete4
PRE
AI
J
Jianxiong Zhang
B
Bing Guo
X
Xuefeng Ding
D
Dasha Hu
J
Jun Tang
K
Ke Du
C
Chao Tang
Y
Yuming Jiang *
DOI:10.1016/j.asoc.2024.111342delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Actual manufacturing process scheduling in enterprise alliances are multi-task scheduling problems involving dynamic factors, and the competition and conflict for manufacturing resources also exist between multitasks. How to perform adaptive multi-objective scheduling of multi-tasks based on the real-time state of the manufacturing environment becomes critical. Therefore, this paper constructs an adaptive multi-task multiobjective scheduling considering resource competition and conflict among tasks (AMMS-RCCT) model based on the enterprise alliance value net, and adopts a hybrid strategy of parallel+serialto resolve conflicts while reducing the waiting time of tasks. With the objective of optimizing the total manufacturing time and total manufacturing cost, an adaptive multi-objective deep Q network (AMDQN) is proposed to solve the AMMSRCCT model. AMDQN is based on a two-hierarchy deep reinforcement learning architecture containing a front controller deep Q network (C-DQN) and a back actuator deep Q network (A-DQN), which performs hierarchical decision-making on optimization objectives and scheduling rules to achieve compromise between multiple objectives while reducing the complexity for optimal selection scheduling rules. For the two optimization objectives of time and cost, two reward algorithms are proposed by introducing two metrics, the estimated tardiness rate and the estimated overspend rate, which guide the A-DQN to learn and adjust the scheduling rules according to the state changes. Besides, nine composite scheduling rules are designed to adapt to the dynamic manufacturing environment from multiple dimensions such as task urgency and completion rate as well as manufacturing resource utilization and cost. Finally, AMDQN is experimentally compared with the proposed nine composite scheduling rules, scheduling rules in existing research, and other scheduling methods based on reinforcement learning in simulated manufacturing environments with different numbers of tasks, subtasks, and manufacturing cells. The experimental results verify the effectiveness and superiority of AMDQN for multi-objective adaptive scheduling in multi-task scheduling problems.
Keywords:
Vaule net
Multi-task scheduling
Resource competition and conflict
Adaptive multi-objective scheduling
Hierarchical deep reinforcement learning

Journal

Applied Soft Computing cover
Applied Soft Computing
IF:
6.6
Papers:
1.4W
Citations:
4.8W

Organization

S
sichuan university
Scholars:
12.0W
Papers: 7.7W
Citations: 100