返回
Semiconductor final test scheduling with Sarsa(λ, k) algorithm
DOI:10.1016/j.ejor.2011.05.052.png)
摘要
En 中文
Semiconductor test scheduling problem is a variation of reentrant unrelated parallel machine problems considering multiple resource constraints, intricate {product, tester, kit, enabler assembly) eligibility constraints, sequence-dependant setup times, etc. A multi-step reinforcement learning (RL) algorithm called Sarsa(lambda, k) is proposed and applied to deal with the scheduling problem with throughput related objective. Allowing enabler reconfiguration, the production capacity of the test facility is expanded and scheduling optimization is performed at the bottom level. Two forms of Sarsa(lambda, k), i.e. forward view Sarsa(lambda, k) and backward view Sarsa(lambda, k), are constructed and proved equivalent in off-line updating. The upper bound of the error of the action-value function in tabular Sarsak(lambda, k) is provided when solving deterministic problems. In order to apply Sarsa(lambda, k), the scheduling problem is transformed into an RL problem by representing states, constructing actions, the reward function and the function approximator. Sarsa(lambda, k) achieves smaller mean scheduling objective value than the Industrial Method (IM) by 68.59% and 76.89%, respectively for real industrial problems and randomly generated test problems. Computational experiments show that Sarsa(lambda, k) outperforms IM and any individual action constructed with the heuristics derived from the existing heuristics or scheduling rules. (C) 2011 Elsevier B.V. All rights reserved.
Keyword:
Scheduling
Semiconductor
Reinforcement learning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6
论文数:
2.2W
被引数:
6.4W
机构
引用论文
A policy gradient method for semi-Markov decision processes with application to call admission control半马尔可夫决策过程的策略梯度方法及其在呼叫接纳控制中的应用

