arXiv AI

Curriculum Learning with GNN-based Reinforcement Learning for Job Shop Scheduling

The paper investigates curriculum learning for graph neural network-based reinforcement learning applied to the job shop scheduling problem. By training policies on progressively larger instances (from 20×20 up to 30×30), the authors demonstrate that this approach reduces training time and improves performance compared to single-size training. Evaluation on unseen instances from 8×8 to 30×30 shows that curriculum learning lowers the mean optimality gap by about 8–9 percentage points and saves roughly 50 hours of training time at the largest target size.

arXiv AI
Sep 17

Variational Approach for Job Shop Scheduling

The paper introduces Variational Graph-to-Scheduler (VG2S), a framework that applies variational inference to the Job Shop Scheduling Problem (JSSP). By decoupling representation learning from policy optimization using a variational graph encoder and an ELBO-based objective, VG2S improves training stability and robustness to hyperparameter changes. Experiments show that VG2S outperforms state‑of‑the‑art deep reinforcement learning baselines and traditional dispatching rules, especially on large‑scale benchmark instances such as DMU and SWV.

By Seung Heon Oh, Jiwon Baek, Hyunjin Oh, Kiyoung Cho, Heechang Yoon, Jong Hun Woo
arXiv Machine Learning
Sep 23

Deep Reinforcement Learning on Item-Compatibility Graphs for One-Dimensional Bin Packing

The paper introduces a novel end‑to‑end, size‑agnostic graph reinforcement learning framework for the one‑dimensional bin packing problem (1D‑BPP). It models packing as a Markov decision process on an item‑compatibility graph, where a graph neural network actor‑critic policy learns to merge compatible partial bins. Empirical results on the BPPLIB benchmark show that the learned policy reduces the mean optimality gap of a constructive heuristic from 2.66 % to 2.31 %, performs competitively against other learned methods, and outperforms a state‑of‑the‑art learned solver on the hardest benchmark family.

By M. Asl{\i} Ayd{\i}n
arXiv AI
Jul 7

A Sliding-Window-Based Reinforcement Learning for Dynamic Assembly Flow Shop Scheduling with Multi-Product Delivery

arXiv:2607. 02941v1 Announce Type: new Abstract: Multi-product kitting delivery imposes significant challenges for real-time scheduling in hybrid manufacturing systems that integrate processing and assembly, as dynamic order arrivals simultaneously alter supply dependencies and the set of feasible job-machine assignments.

By Junhao Qiu, Jianjun Liu, Ting Liu, Rongjie Liao, Zhantao Li, Qingfu Zhang
arXiv AI
Jul 7

Graph Neural Networks are Heuristics

arXiv:2601. 13465v4 Announce Type: replace Abstract: Graph neural networks are usually treated as auxiliaries for combinatorial optimization: they imitate algorithms, guide search, or supply scores to classical procedures.

By Yimeng Min, Carla P. Gomes
arXiv AI
6d ago

PORL: Pretrained Offline Reinforcement Learning for the Job Shop Scheduling Problem

The paper introduces PORL, a hybrid method that first trains a general scheduling policy through online reinforcement learning in simulation, then fine‑tunes it offline on production data using a KL‑divergence constraint to limit policy drift. PORL is evaluated on Job Shop Scheduling Problem instances with distribution shifts and various data sources, consistently outperforming standalone offline RL and other baselines, especially when offline data quality is low. The results suggest that offline adaptation of pretrained policies can improve industrial scheduling when direct online exploration is impractical.

By Mateo Toro Diz, Jonathan Hoss, Noah Klarmann
Hugging Face Trending Papers
Sep 3

PPO-STGNN: A Proximal Policy Optimization Approach with Spatio-Temporal Graph Neural Networks for DAG Task Scheduling in Cloud-Edge-End Computing

PPO-STGNN is a DAG task‑scheduling algorithm that combines proximal policy optimization with spatio‑temporal graph neural networks to address the NP‑hard scheduling problem in heterogeneous cloud‑edge‑end environments. It extracts features from both the DAG task topology and the physical resource graph, then optimizes the scheduling policy to minimize makespan and schedule length ratio while improving CPU and memory load balancing. A multi‑teacher behavior‑cloning pretraining step accelerates convergence, and experiments demonstrate significant load‑balancing gains with low completion times in dynamic, heterogeneous settings.

arXiv AI
Aug 5

PLAN: Parallel Liquid-Inspired Approximation Network for Efficient Representation Learning in Flexible Job Shop Scheduling

arXiv:2608. 03041v1 Announce Type: cross Abstract: Deep reinforcement learning (DRL) approaches for flexible job shop scheduling (FJSP) heavily rely on attention-centric architectures to achieve state-of-the-art performance.

By Dhivya Dharshini Kannan, Wei Zhang, Jieyi Bi, Yingpeng Du, Tianjun Wei, Jie Zhang, Zuming Liu, Anupam Trivedi
arXiv AI
Jun 11

Generalizing Beyond Suboptimality: Offline Reinforcement Learning Learns Effective Scheduling through Random Solutions

arXiv:2509. 10303v2 Announce Type: replace-cross Abstract: Online reinforcement learning (RL) approaches have demonstrated strong performance on Job Shop Scheduling (JSP) and Flexible JSP (FJSP) problems by learning scheduling policies through direct interaction with simulated environments.

By Jesse van Remmerden, Zaharah Bukhsh, Yingqian Zhang
arXiv AI
Sep 4

PPO-STGNN: A Proximal Policy Optimization Approach with Spatio-Temporal Graph Neural Networks for DAG Task Scheduling in Cloud-Edge-End Computing

PPO-STGNN is a DAG task‑scheduling algorithm that combines proximal policy optimization with spatio‑temporal graph neural networks. It extracts features from both the task topology and the heterogeneous cloud‑edge‑end resource graph, then optimizes scheduling to reduce makespan and schedule length ratio while balancing CPU and memory loads. A multi‑teacher behavior‑cloning pretraining step accelerates convergence, and experiments show significant load‑balancing improvements with low completion times in dynamic, heterogeneous environments.

By Yangshuo Qi, Chenwei Wang, Zihan Shen, Songlin Sun