arXiv:2606. 06201v1 Announce Type: new Abstract: Pharmaceutical supply chains (PSCs) struggle with inventory management (IM) due to unpredictable demand patterns and variable lead times associated with restocking.
By Amandeep Kaur, Gyan Prakash
arXiv:2510. 10057v2 Announce Type: replace Abstract: The three-dimensional bin packing problem (3D-BPP) is widely applied in logistics and warehousing.
By Lei Gao, Shihong Huang, Shengjie Wang, Hong Ma, Feng Zhang, Hengda Bao, Qichang Chen, Weihua Zhou
arXiv:2608. 02343v1 Announce Type: cross Abstract: Many operational problems are constrained sequential decision processes with large, combinatorial action spaces and interdependent feasibility constraints.
By Patrick Helm, Jan-Niklas Doerr, Joren Gijsbrechts, Stefan Minner
arXiv:2607. 22356v1 Announce Type: new Abstract: In recent years, the growing complexity of last-mile pickup operations has increased the need for fast and accurate decision-making on logistics platforms.
By Yida Xu, Zhaofang Mao, Yuheng Miao, Jiaxin Zhang, Yiting Sun
arXiv:2607. 02941v1 Announce Type: new Abstract: Multi-product kitting delivery imposes significant challenges for real-time scheduling in hybrid manufacturing systems that integrate processing and assembly, as dynamic order arrivals simultaneously alter supply dependencies and the set of feasible job-machine assignments.
By Junhao Qiu, Jianjun Liu, Ting Liu, Rongjie Liao, Zhantao Li, Qingfu Zhang
arXiv:2606. 18820v1 Announce Type: cross Abstract: Sequential decision problems often exhibit an asymmetric evolution of information and decision flexibility: as a decision cycle unfolds, the agent receives richer information while feasible actions expire due to operational cutoffs, commitments, or resource constraints.
By Jiaxi Liu, Aiping Yang, Yuhang Yang, Shuqi Zhang, Zewei Dong, Jiangming Yang, Xuebin Chen
The paper introduces proximal residual value functions for two‑timescale decision systems, where a planning layer supplies a continuation‑value function to a real‑time optimizer that allocates resources, with inventory placement as a motivating example. The authors propose an end‑to‑end reinforcement learning method that learns a convex residual added to a strictly convex potential, enabling well‑posed optimization and end‑to‑end differentiation while maintaining an explicit convex objective for real‑time execution. They also provide necessary and sufficient conditions for smooth value functions to produce decisions consistent across planning and execution timescales, and demonstrate a 5.0% reduction in routing and transfer cost in an offline simulation using data from a large e‑commerce retailer.
By Harrison Waldon, Carson Eisenach, Akhil Bagaria, Daniel Russo, Dominique Perrault-Joncas, Alisha Zachariah, Dean Foster
arXiv:2607. 28488v1 Announce Type: cross Abstract: Can supply-chain AI move beyond isolated decision modules toward unified operational planning?
By Yunhao Liang, Xianqi Cao, Pujun Zhang, Yuan Qu, Yongzhi Qi, Ningxuan Kang, Max Z. J. Shen
arXiv:2606. 13682v1 Announce Type: new Abstract: The open shop scheduling problem (OSSP) arises in many industrial and service settings but remains computationally challenging as the number of jobs and machines increases.
By Faezeh Ardali, Mwembezi A. Nyelele, Gerald M. Knapp
arXiv:2607. 16875v1 Announce Type: cross Abstract: We introduce the vehicle routing problem with stochastic demands and outsourcing options (VRP-SDO), in which a logistics service provider partitions customer requests into customers outsourced to a common carrier and customers committed to its fixed fleet.
By Mohsen Dastpak, Fausto Errico, Ola Jabali
arXiv:2606. 13604v1 Announce Type: new Abstract: Dispatch in three-sided marketplaces provides a natural setting for reinforcement learning from world feedback: decisions are evaluated by delayed operational outcomes such as delivery speed, courier utilization, and merchant congestion.
By Haochen Wu, Yi Hou, Shiguang Xie
The paper introduces a novel end‑to‑end, size‑agnostic graph reinforcement learning framework for the one‑dimensional bin packing problem (1D‑BPP). It models packing as a Markov decision process on an item‑compatibility graph, where a graph neural network actor‑critic policy learns to merge compatible partial bins. Empirical results on the BPPLIB benchmark show that the learned policy reduces the mean optimality gap of a constructive heuristic from 2.66 % to 2.31 %, performs competitively against other learned methods, and outperforms a state‑of‑the‑art learned solver on the hardest benchmark family.
By M. Asl{\i} Ayd{\i}n