arXiv:2606. 10705v1 Announce Type: cross Abstract: Reinforcement learning promises to optimize sequential decisions in large-scale systems.
By Yavar Yeganeh, Mahsa Shekari, Nicla Frigerio, Daniele Pagano, Andrea Matta
arXiv:2608. 14122v1 Announce Type: new Abstract: Production scheduling in complex manufacturing environments is challenging when sequence-dependent setup times, stochastic disturbances, and due-date constraints must be addressed simultaneously.
By Arne Kr\"oger, Ralf Buscherm\"ohle, Wilhelm Hasselbring, Henrik Wilbers
The paper proposes a hybrid PID–Deep Reinforcement Learning (DRL) controller for industrial processes, addressing the limitations of traditional PID controllers in complex, non‑linear, multi‑input environments. Using the Industrial Benchmark (IB) to test DRL, the authors develop a multi‑objective reward function and employ a TD3 agent to discover optimal settings for the IB’s ‘Gain’ and ‘Shift’ parameters. These parameters are then fed into a tuned PID controller, yielding a system that combines the optimal performance and efficiency of DRL with the reliability of classical control.
By Zhengyang (Cissy), Gu, Joseph E. Hernandez, John Burtenshaw, Sean Scott, Thomas Cook, Chris Couch
The paper introduces a closed‑loop cyber‑physical system for autonomous model lifecycle management in automotive manufacturing, deployed since 2023. It manages paired physics and reinforcement‑learning models, selecting the best candidate through competitive retraining cycles and a Conductor orchestrator that handles plant‑wide inventories and fallback controls. The system incorporates an operator‑trust gate that rejects 23% of policies that deviate from established practice, achieving 28‑45% process stability improvements with no safety incidents.
By Zhengyang (Cissy), Gu, Thomas Cook, Fredaljohn Rohrbaugh, Joseph E. Hernandez, Chris Couch
arXiv:2604.16804v4 Announce Type: replace-cross
Abstract: Optimization problems are central to decision-making in manufacturing, logistics, scheduling, and other industrial settings. Translating comp...
By Sumeet Ramesh Motwani, Chuan Du, Aleksander Petrov, Christopher Davis, Philip Torr, Antonio Papania-Davis, Weishi Yan
AutoOR is a scalable synthetic data generation and reinforcement learning pipeline that trains large language models to autoformalize operations research problems expressed in natural language across linear, mixed‑integer, and non‑linear categories. By generating verified training data from standard optimization forms and using solver execution feedback as a reward signal, AutoOR enables post‑training of an 8B model to achieve state‑of‑the‑art or competitive results on six established OR benchmarks, matching significantly larger frontier models. For non‑linear problems involving physical dynamics, a curriculum RL strategy bootstraps from limited initial data, making this class tractable for post‑training.
By Sumeet Ramesh Motwani, Chuan Du, Aleksander Petrov, Christopher Davis, Philip Torr, Antonio Papania-Davis, Weishi Yan