Event-Driven Reinforcement Learning Enables Long-Horizon Control in Semiconductor Fabrication
arXiv:2606. 10705v1 Announce Type: cross Abstract: Reinforcement learning promises to optimize sequential decisions in large-scale systems.
arXiv:2608. 14122v1 Announce Type: new Abstract: Production scheduling in complex manufacturing environments is challenging when sequence-dependent setup times, stochastic disturbances, and due-date constraints must be addressed simultaneously.
arXiv:2606. 10705v1 Announce Type: cross Abstract: Reinforcement learning promises to optimize sequential decisions in large-scale systems.
The paper introduces PORL, a hybrid method that first trains a general scheduling policy through online reinforcement learning in simulation, then fine‑tunes it offline on production data using a KL‑divergence constraint to limit policy drift. PORL is evaluated on Job Shop Scheduling Problem instances with distribution shifts and various data sources, consistently outperforming standalone offline RL and other baselines, especially when offline data quality is low. The results suggest that offline adaptation of pretrained policies can improve industrial scheduling when direct online exploration is impractical.
arXiv:2607. 02941v1 Announce Type: new Abstract: Multi-product kitting delivery imposes significant challenges for real-time scheduling in hybrid manufacturing systems that integrate processing and assembly, as dynamic order arrivals simultaneously alter supply dependencies and the set of feasible job-machine assignments.
The paper introduces a data‑driven self‑learning control method for highly flexible, modular manufacturing systems. It uses a model‑based reinforcement learning framework that incorporates approximate inverse process models, separating actuation dynamics from state‑space dynamics so that training occurs only in task space. A lightweight feedforward architecture for these inverse models is integrated into standard RL policy networks and tested on a laboratory modular production testbed, showing improved performance and faster training, especially for off‑policy algorithms.
The paper introduces Variational Graph-to-Scheduler (VG2S), a framework that applies variational inference to the Job Shop Scheduling Problem (JSSP). By decoupling representation learning from policy optimization using a variational graph encoder and an ELBO-based objective, VG2S improves training stability and robustness to hyperparameter changes. Experiments show that VG2S outperforms state‑of‑the‑art deep reinforcement learning baselines and traditional dispatching rules, especially on large‑scale benchmark instances such as DMU and SWV.
arXiv:2605. 31044v2 Announce Type: replace Abstract: Reinforcement learning has shown promising results for optimizing the control of industrial energy systems, yet most existing studies remain limited to the application in simulation environments.
The paper introduces a closed‑loop cyber‑physical system for autonomous model lifecycle management in automotive manufacturing, deployed since 2023. It manages paired physics and reinforcement‑learning models, selecting the best candidate through competitive retraining cycles and a Conductor orchestrator that handles plant‑wide inventories and fallback controls. The system incorporates an operator‑trust gate that rejects 23% of policies that deviate from established practice, achieving 28‑45% process stability improvements with no safety incidents.
arXiv:2607. 11725v1 Announce Type: cross Abstract: Prefabricated prefinished volumetric construction moves most building work into module factories, whose production floor operates as a flexible job shop.
arXiv:2606. 06201v1 Announce Type: new Abstract: Pharmaceutical supply chains (PSCs) struggle with inventory management (IM) due to unpredictable demand patterns and variable lead times associated with restocking.
arXiv:2509. 10303v2 Announce Type: replace-cross Abstract: Online reinforcement learning (RL) approaches have demonstrated strong performance on Job Shop Scheduling (JSP) and Flexible JSP (FJSP) problems by learning scheduling policies through direct interaction with simulated environments.
The paper proposes a hybrid PID–Deep Reinforcement Learning (DRL) controller for industrial processes, addressing the limitations of traditional PID controllers in complex, non‑linear, multi‑input environments. Using the Industrial Benchmark (IB) to test DRL, the authors develop a multi‑objective reward function and employ a TD3 agent to discover optimal settings for the IB’s ‘Gain’ and ‘Shift’ parameters. These parameters are then fed into a tuned PID controller, yielding a system that combines the optimal performance and efficiency of DRL with the reliability of classical control.
arXiv:2609.37065v1 Announce Type: new Abstract: Decision-making under uncertainty often relies on predicted parameters, yet accurate prediction does not necessarily lead to good operational decisions...