Benchmarking Multi-Agent Reinforcement Learning for Stochastic OSAT Scheduling: A Reproducible Computational Study for Semiconductor Operations Management
Authors
Doctor of Philosophy Program, KIMT University (Vietnam)
Article Information
DOI: 10.47772/IJRISS.2026.100800210
Subject Category: Learning Modality
Volume/Issue: 10/8 | Page No: 3058-3066
Publication Timeline
Submitted: 2026-08-19
Accepted: 2026-08-24
Published: 2026-08-31
Abstract
Outsourced Semiconductor Assembly and Test (OSAT) operations require repeated scheduling decisions across interdependent stages under processing-time variability, equipment failures, product-mix changes, and demand uncertainty. This study evaluates whether alternative multi-agent reinforcement-learning (MARL) coordination structures provide measurable short-horizon scheduling benefits when compared in a common stochastic OSAT environment. Four learned policy structures—H-MARL-CTDE, H-MARL-Independent, H-MARL-Value, and Flat-MARL—were evaluated against FIFO, SPT, and EDD. The executable simulator represents Die Attach, Wire Bond, Encapsulation, Final Test, and Quality Control, advances in five-minute steps over six simulated hours, and uses active stochastic arrivals, product sampling, lognormal processing-time variation, and MTBF/MTTR-based failure and repair dynamics. The final matched design comprises 27 variability-demand-failure scenarios, ten replications, seven policies, and 1,890 run-level observations. Policy differences were assessed using matched-block Friedman tests and paired Wilcoxon signed-rank tests with Holm correction. H-MARL-Independent achieved the highest mean throughput (16.4932 units/hour) and the lowest energy per completed unit (3.8375), while H-MARL-CTDE achieved the highest OEE proxy (53.690%) and a mean flow time of 0.5848 h. SPT achieved the lowest flow time (0.5838 h). Global policy differences were significant for throughput, flow time, energy efficiency, and OEE (all p<.001), but not makespan (p=.987). CTDE and Independent significantly outperformed FIFO on the four discriminating KPIs after Holm correction, but neither consistently outperformed SPT. The results therefore support selective, KPI-specific learned-policy benefits rather than universal MARL superiority. The primary contribution is an auditable common benchmark that separates computational evidence from production claims and provides a basis for longer-horizon, multi-seed, factory-calibrated, and hybrid RL-heuristic validation in semiconductor operations management.
Keywords
Operations Management; Artificial Intelligence; Decision-Making
Downloads
References
1. Barto, A. G., & Mahadevan, S. (2003). Recent advances in hierarchical reinforcement learning. Discrete Event Dynamic Systems, 13, 41–77. https://doi.org/10.1023/A:1022140919877 [Google Scholar] [Crossref]
2. Cao, Z., Lin, C., Zhou, M., & Huang, R. (2019). Scheduling semiconductor testing facility by using cuckoo search algorithm with reinforcement learning and surrogate modeling. IEEE Transactions on Automation Science and Engineering, 16(2), 825–837. https://doi.org/10.1109/TASE.2018.2862380 [Google Scholar] [Crossref]
3. Chiu, C.-C., Lai, C.-M., Liao, Y.-S., & Yeh, W.-C. (2026). Integration of deep reinforcement learning with simulation optimization applied to semiconductor backend assembly scheduling problem. Swarm and Evolutionary Computation, 100, 102252. https://doi.org/10.1016/j.swevo.2025.102252 [Google Scholar] [Crossref]
4. Dulac-Arnold, G., Levine, N., Mankowitz, D. J., Li, J., Paduraru, C., Gowal, S., & Hester, T. (2021). Challenges of real-world reinforcement learning: Definitions, benchmarks and analysis. Machine Learning, 110, 2419–2468. https://doi.org/10.1007/s10994-021-05961-4 [Google Scholar] [Crossref]
5. Friedman, M. (1937). The use of ranks to avoid the assumption of normality implicit in the analysis of variance. Journal of the American Statistical Association, 32(200), 675–701. https://doi.org/10.1080/01621459.1937.10503522 [Google Scholar] [Crossref]
6. Ghavamzadeh, M., Mahadevan, S., & Makar, R. (2006). Hierarchical multi-agent reinforcement learning. Autonomous Agents and Multi-Agent Systems, 13(2), 197–229. https://doi.org/10.1007/s10458-006-7035-4 [Google Scholar] [Crossref]
7. Holm, S. (1979). A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6(2), 65–70. [Google Scholar] [Crossref]
8. Lin, J., Li, Y.-Y., & Song, H.-B. (2022). Semiconductor final testing scheduling using Q-learning based hyper-heuristic. Expert Systems with Applications, 187, 115978. https://doi.org/10.1016/j.eswa.2021.115978 [Google Scholar] [Crossref]
9. Lowe, R., Wu, Y., Tamar, A., Harb, J., Abbeel, P., & Mordatch, I. (2017). Multi-agent actor-critic for mixed cooperative-competitive environments. Advances in Neural Information Processing Systems, 30, 6379–6390. [Google Scholar] [Crossref]
10. Mayerhoff, J., & Schmidt, M. (2026). Reinforcement learning for autonomous production planning and control: A systematic literature review. Journal of Manufacturing Systems, 86, 546–568. https://doi.org/10.1016/j.jmsy.2026.03.023 [Google Scholar] [Crossref]
11. Panzer, M., & Bender, B. (2022). Deep reinforcement learning in production systems: A systematic literature review. International Journal of Production Research, 60(13), 4316–4341. https://doi.org/10.1080/00207543.2021.1973138 [Google Scholar] [Crossref]
12. Park, I. B., & Park, J. (2023). Scalable scheduling of semiconductor packaging facilities using deep reinforcement learning. IEEE Transactions on Cybernetics, 53(6), 3518–3531. https://doi.org/10.1109/TCYB.2021.3128075 [Google Scholar] [Crossref]
13. Pinedo, M. L. (2022). Scheduling: Theory, algorithms, and systems (6th ed.). Springer. https://doi.org/10.1007/978-3-031-05921-6 [Google Scholar] [Crossref]
14. Rashid, T., Samvelyan, M., Schroeder de Witt, C., Farquhar, G., Foerster, J., & Whiteson, S. (2018). QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning. Proceedings of the 35th International Conference on Machine Learning, 80, 4295–4304. https://proceedings.mlr.press/v80/rashid18a.html [Google Scholar] [Crossref]
15. Schulman, J., Moritz, P., Levine, S., Jordan, M. I., & Abbeel, P. (2016). High-dimensional continuous control using generalized advantage estimation. International Conference on Learning Representations. https://arxiv.org/abs/1506.02438 [Google Scholar] [Crossref]
16. Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv. https://doi.org/10.48550/arXiv.1707.06347 [Google Scholar] [Crossref]
17. Stöckermann, P., Feudel, S., Immordino, A., Hayen, N., Altenmüller, T., Gebser, M., Schekotihin, K., Seidel, G., Wegmann, M., & Higgins, F. (2025). Reinforcement learning based dispatching solutions in semiconductor manufacturing: A literature review on validation and deployment. Production & Manufacturing Research, 13(1), 2582472. https://doi.org/10.1080/21693277.2025.2582472 [Google Scholar] [Crossref]
18. Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W. M., Zambaldi, V., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J. Z., Tuyls, K., & Graepel, T. (2018). Value-Decomposition Networks for cooperative multi-agent learning based on team reward. Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems, 2085–2087. [Google Scholar] [Crossref]
19. Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press. [Google Scholar] [Crossref]
20. Wilcoxon, F. (1945). Individual comparisons by ranking methods. Biometrics Bulletin, 1(6), 80–83. https://doi.org/10.2307/3001968 [Google Scholar] [Crossref]
21. Zhang, Z., Zheng, L., Hou, F., & Li, N. (2011). Semiconductor final test scheduling with Sarsa(λ, k) algorithm. European Journal of Operational Research, 215(2), 446–458. https://doi.org/10.1016/j.ejor.2011.05.052 [Google Scholar] [Crossref]
22. Zhang, L., Lin, Y., Xu, C., & Liu, M. (2024). A new EDA algorithm combined with Q-learning for semiconductor final testing scheduling problem. Computers & Industrial Engineering, 193, 110259. https://doi.org/10.1016/j.cie.2024.110259 [Google Scholar] [Crossref]