arXiv AI

Event-Driven Reinforcement Learning Enables Long-Horizon Control in Semiconductor Fabrication

arXiv:2606. 10705v1 Announce Type: cross Abstract: Reinforcement learning promises to optimize sequential decisions in large-scale systems.

arXiv AI
Aug 17

Reinforcement Learning-Based Production Scheduling in an Industry-Based Coating Scenario Using the Digital Model Playground

arXiv:2608. 14122v1 Announce Type: new Abstract: Production scheduling in complex manufacturing environments is challenging when sequence-dependent setup times, stochastic disturbances, and due-date constraints must be addressed simultaneously.

By Arne Kr\"oger, Ralf Buscherm\"ohle, Wilhelm Hasselbring, Henrik Wilbers
arXiv Machine Learning
Sep 11

Distributed Optimization of Modular Production Systems using Model-based Reinforcement Learning with Inverse Models

The paper introduces a data‑driven self‑learning control method for highly flexible, modular manufacturing systems. It uses a model‑based reinforcement learning framework that incorporates approximate inverse process models, separating actuation dynamics from state‑space dynamics so that training occurs only in task space. A lightweight feedforward architecture for these inverse models is integrated into standard RL policy networks and tested on a laboratory modular production testbed, showing improved performance and faster training, especially for off‑policy algorithms.

By Andreas Schwung, Steve Yuwono, Sofiene Lassoued, Dorothea Schwung
arXiv Machine Learning
Sep 22

Autonomous Model Lifecycle Management for Digital Twin-Based Manufacturing Control

The paper introduces a closed‑loop cyber‑physical system for autonomous model lifecycle management in automotive manufacturing, deployed since 2023. It manages paired physics and reinforcement‑learning models, selecting the best candidate through competitive retraining cycles and a Conductor orchestrator that handles plant‑wide inventories and fallback controls. The system incorporates an operator‑trust gate that rejects 23% of policies that deviate from established practice, achieving 28‑45% process stability improvements with no safety incidents.

By Zhengyang (Cissy), Gu, Thomas Cook, Fredaljohn Rohrbaugh, Joseph E. Hernandez, Chris Couch
arXiv Machine Learning
Sep 22

Augmenting PID Control with Deep Reinforcement Learning: A Hybrid Approach to the Industrial Benchmark

The paper proposes a hybrid PID–Deep Reinforcement Learning (DRL) controller for industrial processes, addressing the limitations of traditional PID controllers in complex, non‑linear, multi‑input environments. Using the Industrial Benchmark (IB) to test DRL, the authors develop a multi‑objective reward function and employ a TD3 agent to discover optimal settings for the IB’s ‘Gain’ and ‘Shift’ parameters. These parameters are then fed into a tuned PID controller, yielding a system that combines the optimal performance and efficiency of DRL with the reliability of classical control.

By Zhengyang (Cissy), Gu, Joseph E. Hernandez, John Burtenshaw, Sean Scott, Thomas Cook, Chris Couch
arXiv AI
Jul 7

A Sliding-Window-Based Reinforcement Learning for Dynamic Assembly Flow Shop Scheduling with Multi-Product Delivery

arXiv:2607. 02941v1 Announce Type: new Abstract: Multi-product kitting delivery imposes significant challenges for real-time scheduling in hybrid manufacturing systems that integrate processing and assembly, as dynamic order arrivals simultaneously alter supply dependencies and the set of feasible job-machine assignments.

By Junhao Qiu, Jianjun Liu, Ting Liu, Rongjie Liao, Zhantao Li, Qingfu Zhang
arXiv Machine Learning
Sep 22

Proximal Residual Value Functions for Consistent Planning and Real-Time Execution

The paper introduces proximal residual value functions for two‑timescale decision systems, where a planning layer supplies a continuation‑value function to a real‑time optimizer that allocates resources, with inventory placement as a motivating example. The authors propose an end‑to‑end reinforcement learning method that learns a convex residual added to a strictly convex potential, enabling well‑posed optimization and end‑to‑end differentiation while maintaining an explicit convex objective for real‑time execution. They also provide necessary and sufficient conditions for smooth value functions to produce decisions consistent across planning and execution timescales, and demonstrate a 5.0% reduction in routing and transfer cost in an offline simulation using data from a large e‑commerce retailer.

By Harrison Waldon, Carson Eisenach, Akhil Bagaria, Daniel Russo, Dominique Perrault-Joncas, Alisha Zachariah, Dean Foster