arXiv Machine Learning

OR-Transformer: Scaling Real-Time Decision-Making to 1,000 Items

OR-Transformer is a deep reinforcement learning framework designed for joint replenishment in supply chain operations with thousands of items. It uses a permutation‑equivariant Transformer architecture and pathwise‑gradient training to handle high‑dimensional observation and action spaces. In tests up to 1,024 items, it outperforms both learning‑based and rolling‑horizon MILP baselines and cuts online decision time by over four million times.

arXiv AI
Jul 7

A Sliding-Window-Based Reinforcement Learning for Dynamic Assembly Flow Shop Scheduling with Multi-Product Delivery

arXiv:2607. 02941v1 Announce Type: new Abstract: Multi-product kitting delivery imposes significant challenges for real-time scheduling in hybrid manufacturing systems that integrate processing and assembly, as dynamic order arrivals simultaneously alter supply dependencies and the set of feasible job-machine assignments.

By Junhao Qiu, Jianjun Liu, Ting Liu, Rongjie Liao, Zhantao Li, Qingfu Zhang
arXiv AI
Jun 18

Maturing Markov Decision Processes: Decision Making under Increasing Information and Shrinking Action Sets

arXiv:2606. 18820v1 Announce Type: cross Abstract: Sequential decision problems often exhibit an asymmetric evolution of information and decision flexibility: as a decision cycle unfolds, the agent receives richer information while feasible actions expire due to operational cutoffs, commitments, or resource constraints.

By Jiaxi Liu, Aiping Yang, Yuhang Yang, Shuqi Zhang, Zewei Dong, Jiangming Yang, Xuebin Chen
arXiv Machine Learning
Sep 22

Proximal Residual Value Functions for Consistent Planning and Real-Time Execution

The paper introduces proximal residual value functions for two‑timescale decision systems, where a planning layer supplies a continuation‑value function to a real‑time optimizer that allocates resources, with inventory placement as a motivating example. The authors propose an end‑to‑end reinforcement learning method that learns a convex residual added to a strictly convex potential, enabling well‑posed optimization and end‑to‑end differentiation while maintaining an explicit convex objective for real‑time execution. They also provide necessary and sufficient conditions for smooth value functions to produce decisions consistent across planning and execution timescales, and demonstrate a 5.0% reduction in routing and transfer cost in an offline simulation using data from a large e‑commerce retailer.

By Harrison Waldon, Carson Eisenach, Akhil Bagaria, Daniel Russo, Dominique Perrault-Joncas, Alisha Zachariah, Dean Foster
arXiv AI
Jul 21

A Deep Reinforcement Learning Algorithm for the Vehicle Routing Problem with Stochastic Demands and Outsourcing

arXiv:2607. 16875v1 Announce Type: cross Abstract: We introduce the vehicle routing problem with stochastic demands and outsourcing options (VRP-SDO), in which a logistics service provider partitions customer requests into customers outsourced to a common carrier and customers committed to its fixed fleet.

By Mohsen Dastpak, Fausto Errico, Ola Jabali
arXiv AI
Jun 12

Multi-Agent Reinforcement Learning from Delayed Marketplace Feedback for Objective-Weight Adaptation in Three-Sided Dispatch

arXiv:2606. 13604v1 Announce Type: new Abstract: Dispatch in three-sided marketplaces provides a natural setting for reinforcement learning from world feedback: decisions are evaluated by delayed operational outcomes such as delivery speed, courier utilization, and merchant congestion.

By Haochen Wu, Yi Hou, Shiguang Xie
arXiv Machine Learning
Sep 23

Deep Reinforcement Learning on Item-Compatibility Graphs for One-Dimensional Bin Packing

The paper introduces a novel end‑to‑end, size‑agnostic graph reinforcement learning framework for the one‑dimensional bin packing problem (1D‑BPP). It models packing as a Markov decision process on an item‑compatibility graph, where a graph neural network actor‑critic policy learns to merge compatible partial bins. Empirical results on the BPPLIB benchmark show that the learned policy reduces the mean optimality gap of a constructive heuristic from 2.66 % to 2.31 %, performs competitively against other learned methods, and outperforms a state‑of‑the‑art learned solver on the hardest benchmark family.

By M. Asl{\i} Ayd{\i}n