arXiv:2606. 10705v1 Announce Type: cross Abstract: Reinforcement learning promises to optimize sequential decisions in large-scale systems.
By Yavar Yeganeh, Mahsa Shekari, Nicla Frigerio, Daniele Pagano, Andrea Matta
The paper introduces PORL, a hybrid method that first trains a general scheduling policy through online reinforcement learning in simulation, then fine‑tunes it offline on production data using a KL‑divergence constraint to limit policy drift. PORL is evaluated on Job Shop Scheduling Problem instances with distribution shifts and various data sources, consistently outperforming standalone offline RL and other baselines, especially when offline data quality is low. The results suggest that offline adaptation of pretrained policies can improve industrial scheduling when direct online exploration is impractical.
By Mateo Toro Diz, Jonathan Hoss, Noah Klarmann
arXiv:2607. 02941v1 Announce Type: new Abstract: Multi-product kitting delivery imposes significant challenges for real-time scheduling in hybrid manufacturing systems that integrate processing and assembly, as dynamic order arrivals simultaneously alter supply dependencies and the set of feasible job-machine assignments.
By Junhao Qiu, Jianjun Liu, Ting Liu, Rongjie Liao, Zhantao Li, Qingfu Zhang
The paper introduces a data‑driven self‑learning control method for highly flexible, modular manufacturing systems. It uses a model‑based reinforcement learning framework that incorporates approximate inverse process models, separating actuation dynamics from state‑space dynamics so that training occurs only in task space. A lightweight feedforward architecture for these inverse models is integrated into standard RL policy networks and tested on a laboratory modular production testbed, showing improved performance and faster training, especially for off‑policy algorithms.
By Andreas Schwung, Steve Yuwono, Sofiene Lassoued, Dorothea Schwung
The paper introduces Variational Graph-to-Scheduler (VG2S), a framework that applies variational inference to the Job Shop Scheduling Problem (JSSP). By decoupling representation learning from policy optimization using a variational graph encoder and an ELBO-based objective, VG2S improves training stability and robustness to hyperparameter changes. Experiments show that VG2S outperforms state‑of‑the‑art deep reinforcement learning baselines and traditional dispatching rules, especially on large‑scale benchmark instances such as DMU and SWV.
By Seung Heon Oh, Jiwon Baek, Hyunjin Oh, Kiyoung Cho, Heechang Yoon, Jong Hun Woo
arXiv:2605. 31044v2 Announce Type: replace Abstract: Reinforcement learning has shown promising results for optimizing the control of industrial energy systems, yet most existing studies remain limited to the application in simulation environments.
By Tobias Lademann, Th\'eo Vincent, Jan Peters, Matthias Weigold